Biomedical and toxicological research has long relied heavily on animal models and simple cell cultures. Yet their limitations, particularly in predicting human responses, have contributed to the emergence of New Approach Methodologies (NAMs): human-based and computational approaches intended to improve scientific relevance while reducing animal use.
Among these NAMs, organoids, organs-on-chips and other advanced in vitro models are especially important as they can accurately replicate human physiology, including tissue structure, flow, mechanical forces and interactions between cell types, more accurately than conventional methods.[1] Their significance lies not only in generating more data, but also in producing evidence that closely reflect human conditions that researchers and regulators need to understand.
This article explores a question that arises from this progress: once advanced in vitro models produce richer evidence, how can researchers interpret them reliably? Here, we focus on discussing the role of artificial intelligence in identifying patterns, integrating measurements and guiding experimental decisions, while recognizing that biological validation and a clearly defined context of use remain essential.
When In Vitro Models Outpace Our Ability to Interpret Them
Beyond images and molecular profiles, these models also yield dose responses, donor-specific differences, sensor traces and repeated measurements over time. This richness allows biological responses to be observed in greater detail, but it also introduces a new analytical challenge. As data generation becomes increasingly tractable, interpretation emerges as the limiting step. Researchers must determine how these different measurements relate to one another and what they collectively represent for the question being studied.
The difficulty has several dimensions: First, the data are high-dimensional, with many variables changing simultaneously. Second, they are also heterogeneous, as images, molecular profiles and functional sensors each describe biology in different ways. Third, they span multiple scales, from molecular and cellular events to tissue-level behavior. Lastly, they are dynamic, as processes such as injury, adaptation and recovery unfold over time. Therefore, no single readout is sufficient to capture the full physiological state. A tissue may appear structurally intact while its metabolism changes, and a molecular average may obscure a small but biologically significant subpopulation of damaged cells.
The practical task is therefore to identify meaningful signals, connect evidence across data types and translate the result into a defensible next step. AI can assist by organizing complex observations, revealing patterns worth investigating and indicating where direct measurement is most needed. Its value lies not simply in processing more data, but in helping researchers interpret them without losing sight of their biological context, uncertainty and limitations that surround every measurement.

What AI Can Add to In Vitro Interpretation
First, AI can recognize patterns distributed across many measurements simultaneously. A human observer readily notices widespread cell death or a broken barrier, but early stress may appear as subtle, concurrent changes in nuclear shape, organelle texture and cell spacing. No single change is conclusive in isolation, but together they may describe a reproducible state. AI can score these combined signatures consistently across large datasets, revealing which treated samples resemble controls, which donors respond differently, and which conditions are drifting toward abnormality.[2]
Second, AI can integrate different biological languages. Images, gene expression, secreted proteins, metabolism and sensor traces may each reflect the same event from a different perspective.[3] Aligning these layers can reveal agreement across measurements but disagreement is equally informative. Cells may appear structurally intact under the microscope while their secreted signals or metabolism already register stress. Neither observation alone tells the whole story. The purpose of integration is not to consolidate every measurement into one answer, but to expose relationships that deserve follow-up.
Third, AI can support virtual or non-destructive readouts. A model trained on the relationship between label-free imaging and a direct but destructive assay can estimate the health status of living samples without intervening [4,5] This does not imply the direct assay redundant. Rather, it allows researchers to monitor samples longitudinally, preserve scarce human-derived material and identify the cultures for which direct measurement is most important.
Taken together, these capabilities position AI as an interpretation layer for modern in vitro research: a way to monitor biological change, compare donor responses, integrate evidence across assays and make better use of scarce human material. The next question is how this translates into day-to-day experimental work.
Practical Examples: Monitoring, Screening and Prioritization
Consider a patient-derived organoid after drug exposure. Researchers may want to know whether it is viable, stressed, differentiating, recovering or approaching irreversible injury. Staining, genetic reporters and biochemical assays can provide direct evidence, but some consume or alter the sample. Brightfield imaging is repeatable and label-free, yet subtle damage can be difficult to distinguish from normal structure or matrix artifacts.
In one example, researchers paired brightfield images of organoids grown in a supporting gel with direct viability measurements. The resulting model estimated whether new cultures appeared viable or damaged.[5] Used carefully, this changes the workflow: cultures can be followed over time, while direct assays are reserved for selected conditions or used to confirm uncertain predictions.
AI-assisted imaging can also help researchers anticipate how a culture is likely to develop. In small airway-on-a-chip models, epithelial differentiation is an important sign that the tissue has matured successfully. Researchers have shown that brightfield images collected early in culture can help predict later differentiation.[6] This could identify cultures unlikely to develop as intended before weeks of work and resources have been invested.
The same methodology can also guide screening. In a broader high-throughput imaging context, image-derived features from a large compound screen have been used to predict activity across many biochemical and cellular assays, allowing later projects to test smaller groups of predicted-active compounds.[7] The images did not replace those assays; they changed which compounds entered them.
AI can also help choose the next experiment. Dose, time point, medium, matrix, flow rate, donor and readout can create more combinations than a laboratory can test. Iterative methods learn from early results and suggest conditions that are either promising or especially informative. Similar iterative approaches have been used to optimize cell culture media and to decide which measurements justify their cost.[8,9] Across monitoring, screening and experimental design, the contribution is the same: focus direct measurement where it can resolve the most important uncertainty.
Why AI Can Mislead if the Data Are Weak
These applications are promising, but they hold only when the underlying data are sound. An AI model may learn technical differences rather than biology. Imaging conditions, experimental dates, reagent lots, plate positions and instruments can all influence the data. If treated and control samples are prepared differently, the model may learn these batch effects instead of the treatment response. Correcting too little leaves technical noise, while correcting too much may remove real biological variation.[10]
Label quality matters as well. A model trained to predict “toxicity” learns how toxicity was defined in its training data, whether through viability, morphology or expert judgment. None captures every form of injury. Performance against a narrow label does not prove that the model understands the wider biological response.
Finally, explainability should not be confused with mechanism. A highlighted image region or influential molecular feature may show what contributed to a prediction, but it does not establish why the biological response occurred. AI can identify patterns worth investigating; biological experiments are still needed to explain them.
Together, these limitations show that model performance alone is not enough. Before an AI-interpreted result can support a scientific, clinical or regulatory decision, researchers must define what the model is intended to do, where it should be used and what evidence is needed to validate it.

Trust, Context and Human-Relevant Decisions
The practical question is not whether AI is powerful, but what evidence is needed before an AI-interpreted in vitro result can change a real decision. This is also the direction taken by regulators. The U.S. Food and Drug Administration has proposed a risk-based credibility framework for AI models used to support regulatory decisions, while the European Medicines Agency has issued a reflection paper on AI across the medicinal product lifecycle.[11,12] In 2026, the FDA also accepted the first Letter of Intent for an AI-driven in silico drug development tool into the ISTAND qualification program, to help predict drug-induced liver injury.[13] This was not an approval, but it shows that such tools are beginning to enter formal qualification pathways.
Trust depends on context of use (CoU). A model used to rank compounds in early research can tolerate more uncertainty than one used to dismiss a safety concern or support a regulatory claim. A credible claim must specify the system, data type, endpoint, chemical domain and decision supported. Validation must also match the distance of the prediction: predicting within a familiar dataset is different from predicting a new donor, laboratory or human exposure. For this reason, validation should test genuinely new donors, experimental batches or laboratories where possible, rather than relying only on randomly held-out samples from the same dataset. Useful models should also express uncertainty, flag unfamiliar samples and request direct measurement when evidence is insufficient.
This also clarifies the relationship between AI and animal replacement. On its own, an AI prediction does not yet make an in vitro result a reliable alternative to animal studies. Its contribution is strongest when it improves the credibility of human-relevant approaches: strengthening quality control, extracting more information from scarce human material, linking in vitro measurements to clinical questions and guiding experiments toward the uncertainties that matter most. With clearer contexts, stronger validation and repeated testing in real use, AI-assisted NAMs may increasingly support decisions that have historically relied on animals. The goal is to use AI to strengthen the interpretation of human-relevant biology, while keeping biological evidence at the center.