In spatial biology, context is not metadata added after the experiment. Context is part of the measurement itself.
Consider two lung cancer tissue sections analyzed using spatial transcriptomics. One was collected from a primary tumor before treatment, while the other came from a metastatic lesion after multiple lines of therapy. Even if the same genes and cells were technically measured, the implications of the two results could be entirely different.
Yet in real-world datasets, these differences are often confined to a small table labeled “sample information” or omitted altogether. The number of cells may have increased, but if we do not know which patient or clinical time point those cells represent, data volume and interpretability do not increase together.
The critical scarce resource in spatial transcriptomics is not just the equipment. It is the right tissue and the clinical context needed to describe it.
A Larger Dataset Is Not Necessarily a Better Dataset
Dataset size is often expressed in terms of the number of patients, tissue samples, or cells. These figures indicate throughput, but they do not guarantee that the data are representative or suitable for drug development.
Even within the same cancer type, biology can vary according to disease stage, histological subtype, treatment history, ethnic background, and collection site. Primary tumors differ from metastatic lesions, and the tumor microenvironment may vary depending on the organ to which the cancer has metastasized. Differences in preprocessing and quality also exist between surgical and biopsy specimens, as well as between fresh-frozen and formalin-fixed, paraffin-embedded tissues.
A dataset containing 1,000 patients may therefore have limited applicability if it is skewed toward a narrow set of indications or clinical settings. Conversely, a cohort of only a few dozen patients may be more useful for critical decision-making if it has been carefully designed around a specific clinical question.
At the surface, the challenge lies in obtaining tissue. Beneath it is a complex interplay of patient consent, pathology quality, linkage between clinical information and tissue, institution-specific storage practices, preprocessing time, and data usage rights. Ultimately, the tissue bottleneck is not simply a shortage of material. It is a shortage of tissue that can be used, adequately described, and meaningfully compared.
Tumor-Only Data Make It Difficult to Assess the Therapeutic Window
When evaluating a drug target, high expression in tumors alone is not enough. Expression levels and spatial distribution in normal tissues must also be examined to estimate the range between efficacy and toxicity, or the therapeutic window.
Tumor-only datasets are useful for understanding intratumoral heterogeneity and the tumor microenvironment, but they cannot reveal which cells in normal organs such as the kidney, liver, heart, and gastrointestinal tract express the target. Adjacent normal tissue is also informative, but because it may be affected by inflammation or field effects, it cannot fully substitute for healthy normal organs.
This is why tumor and normal data need to be built in parallel. Ideally, tumor and normal tissues from the same patient should be linked. Depending on the objective, however, comparison with an independent normal-organ reference may be sufficient. Because obtaining perfectly matched samples across every cancer type is not realistic, priority should be given to the normal tissues representing the key organs at risk for toxicity from the relevant therapeutic modality.
Balancing Standardization and Diversity
Single-institution cohorts offer the advantage of relatively consistent pathology review, clinical information, and tissue-processing methods. However, they may be biased toward a particular region or patient population. Multi-institutional cohorts can improve generalizability, but they must accommodate variation arising from differences in surgical, storage, and preprocessing procedures.
The solution is not to force every tissue sample into a single standardized process.
In the short term, cohorts should first be designed around the research objective (external link: http://link.portrai.io or the Cohortlab overview page), and minimum clinical metadata should be defined, including disease stage, treatment history, collection site, and whether the sample is primary or metastatic. Pathology QC and analytical QC should not be treated separately; the effects of tissue condition on data quality should also be documented.
In the medium term, tumor-normal references should be established for major indications, followed by the gradual expansion of data across multiple regions and institutions. In the long term, hospitals, data companies, and pharmaceutical companies should form sustained data partnerships rather than relying on one-off specimen purchases. The value of data accumulates when tissue and clinical outcomes remain connected over time.
However, for inherently difficult-to-obtain tissues, such as those from rare cancers or longitudinal samples collected before and after treatment, demanding large-scale standardization may exclude precisely the data that matter most. In such cases, evaluation should focus less on scale and more on the rarity of the samples, the depth of the clinical information, and the ability to track them over time.
What Portrai Has Learned
In building human-derived spatial datasets across multiple cancer types, Portrai has learned that the value of data cannot be explained by cell count alone. The meaning of target expression varied depending on the patient and anatomical site from which it was measured, whether the sample was collected before or after treatment, and whether the expression occurred in cancer cells or normal cells.
For modalities such as ADCs, RPTs, and TCEs in particular, it was necessary to consider not only intratumoral expression levels but also expression in normal organs, consistency across patients, and proximity to surrounding cells. This experience has led us to ask “What decision will this cohort be used to inform?” before competing on dataset size.
The Next Five Years
In an optimistic scenario, multi-regional and multi-ethnic spatial cohorts will be linked to clinical outcomes, enabling the generalizability of targets and biomarkers to be assessed earlier. More realistically, purpose-built cohorts are likely to accumulate first in major cancer types and for high-value therapeutic modalities.
Conversely, if large-scale datasets are produced without clinical context, future AI models may be technically sophisticated yet fail to adequately represent real patient populations. A good model requires not only large datasets, but also an understanding of who the data came from, when they were generated, and why.
In spatial biology, context is not metadata added after the experiment. Context is part of the measurement itself.
Securing high-quality tissue alone, however, is not enough. To process limited clinical samples at scale and with consistent quality, the next bottleneck must be addressed: the standardization and automation of experiments.










