Skip to content
Back to Insights
portrai-insight

[Sid’s Data Journey - Part 3] Virtual Tissue Pharmacology: Bringing In Silico Drug Testing to a Patient Tissue

The AI scientist in the loop Two things happen in this part. We turn expression into radiation dose, which is the number aradiopharmaceutical is actually judged on. And we show who did the first vers…

HC
Hongyoon Choi - 13 min read
Share

The AI scientist in the loop

Two things happen in this part. We turn expression into radiation dose, which is the number aradiopharmaceutical is actually judged on. And we show who did the first version of thatcalculation, which was not a person.

Dose, at last

Everything in the previous part is expression. A radioligand or ADC does not deliver expression. It delivers energy, measured in grays (for radioligand), and how much arrives depends on thingsexpression cannot tell you: how the drug escapes the blood vessels, how far it diffuses throughtissue, how tightly it binds, how fast the cell swallows it, and how far the radiation travels afterthe atom decays.

We ran that simulation (Portrai Microscopic-PK) for the top three candidates plus both referencetargets, on both tissue cores, with 7,400 MBq of Lu-177 and a physical half-life of 6.7 days. Lu-177 beta particles travel, so the dose in any square depends on its neighbours as well as itself.

p01_img001_674x472.png

One shared colour scale across all five panels, logarithmic because the dose distribution spansfive orders of magnitude. Both cores were simulated, and we report them separately rather thanaveraged, because the average hides a disagreement.

09_six_axis_summary.jpgp02_img002_760x445.png

Three things in that table, in decreasing order of how much we believe them.

FAP is last on both cores, by both measures. Roughly half the mean dose of the leaders, and 3to 13 times less of the tumour clearing 1 Gy. This is the one conclusion the two cores agree oncompletely, and it is the one that matters.

TNC and B7-H3 swap places between the two cores. TNC wins on B3, B7-H3 wins on C3, andneither margin is large. Two cores from one patient are not enough to separate the top twocandidates, and we are not going to pretend otherwise by quoting the average. What survives isthat both are clearly ahead of FAP, and ahead of ITGA10 and PTK7.

p03_img003_536x333.pngp03_img004_638x433.png

That second figure makes the mechanism visible. FAP actually delivers the highest mediandose of the five, around 0.27 Gy against 0.05 to 0.16 for the others, and then falls off a cliff. Theothers deliver less to the typical square and far more to the squares that matter.

What is happening is that FAP's few positive squares are so far apart that their radiation smearsout into a thin uniform wash instead of building up anywhere. Crossfire is adding dose frommany distant weak sources rather than concentrating it. That is the direct consequence of thegeometry in the previous section: the average gap from a dim square to a bright one was 350micrometres for FAP and 111 for B7-H3, and a Lu-177 beta particle travels about 280.

A uniform 0.3 Gy wash does not kill a tumour. Concentrated dose does.

Two caveats on all of the above. The numbers are in grays but they are not calibrated: the modelmaps a normalised expression value onto a micromolar receptor concentration through a singlescaling factor, so read 8 Gy as "about twice 3.5 Gy" and NOT as an ‘absolute’clinical doseprediction. And the simulation runs 10 hours, which covers uptake and early clearance but not

the multi-day tail of a real treatment cycle. The comparison between targets is the result. Theabsolute grays are not.

Who ran this first

The dose simulation above was not the first one performed on this patient. Six months beforethis study, someone typed a single line into this project's chat window:

@Sid-T0-B3-D001 FAP targeting RPT, FAP-2286-based kineticprofiles

What came back was not a result. It was a proposal. Here is that exchange, still sitting in theapplication today.

p04_img005_760x452.png

One line of input on the left, and a seventeen-row parameter table in reply. Look at what is in thethird row of that table: Fibroblast Activation Protein (46/21,040 spots express). Before proposinganything the agent went and counted, and it put the count in the proposal where the personapproving it would see it.

p05_img006_603x696.png

Almost entirely black: the dose is near zero nearly everywhere. It is the same conclusion this study reached by a completely different route.

What the agent actually did 

The project database still holds the whole exchange: 3 sessions, 18 messages, and 98 recorded reasoning steps broken down as 41 iterations, 20 actions, 19 observations, 8 final answers, 6 plans and 4 explicit thoughts. 40 output files are registered against them. 

The sequence for the FAP run went like this. 

It looked at the data before proposing anything. Its first recorded observation: 

Sample: Sid-T0-B3-D001 Spots: 21040, Genes: 17644

FAP found! Expression range: [0.000, 7.132], mean=0.011

Spots with FAP>0: 46 / 21040

Endothelial markers available: ['VWF', 'PECAM1', 'ENG', 'CDH5']

46 spots out of 21,040. It found the central fact of this whole series in its first tool call, and it found it because it checked whether the target existed before agreeing to simulate it. 

It proposed parameters and justified each one. A seventeen-row table: Lu-177 because that iswhat FAP-2286 uses clinically, 7,400 MBq because that is a standard therapeutic activity, adose point kernel for beta crossfire. And a clearance rate it argued for rather than defaulted to:

FAP-2286 is a peptide-based FAP inhibitor (not an antibody), so it has muchfaster clearance than antibody-based RPTs, hence half_clear=0.5 h (vs~10 h for antibodies).

That is the right distinction and it is not a lookup. It is the kind of thing a reviewer would askabout.

It stopped and waited. The message ended:

Please confirm these parameters, or tell me what to adjust.

A human typed Go, seven minutes later. Only then did it run anything, and when it was done itreported back into the same thread.

p06_img007_760x452.png

Sample, spot count, target with its expression range, isotope, dose, simulation length, then a listof what it wrote to disk. The last row of that list is the one the next section uses: a time-seriesviewer layer with a slider.

p07_img008_760x475.png

The simulation is not a picture, it is twenty of them

That curve is the whole tissue averaged into one line per timepoint. The simulation the agentsaved keeps all 21,040 bins at all 20 timepoints, and the viewer will play it.

p07_img009_760x452.png

Same tissue, same colour scale, three positions of the time slider, and the viewer's own mean inthe corner each time: 0.009, then 0.060, then 0.044 MBq/mL.

At the moment of injection almost nothing has arrived. What little has is a scatter of isolatedbright bins, which are the ones sitting next to a vessel. Forty-three minutes later the drug isthrough most of the core, and it is through it unevenly: bright along the top edge and the rightmargin, dim in the middle. By ten hours it is washing out and the whole field has dimmed.

That middle panel is the useful one. A time-activity curve tells you the tissue average peakedand fell. It cannot tell you that the centre of the core saw less than the rim, and that is a questionabout whether a drug works.

This is what the word "world model" is doing in the project's name. The output is not a numberfor the tumour, it is a concentration at every point in the tissue at every moment, and thequestions you can ask of it are geometric ones: did the drug reach the middle, how long did itstay, which parts never saw it.

The same agent, a different question

The FAP simulation was not the only thing it did. Asked to predict drug sensitivity, it producedper-cell maps across all three biopsies.

p08_img010_760x863.png

Each dot is a 32-micrometre square of the patient's tumour, coloured by predicted sensitivity toifosfamide. And because it ran the same six drugs on all three timepoints, there is a trajectory.

p09_img011_760x432.png11_drug_response_time.jpg

Lower means more predicted sensitivity. Every drug moves the same direction, and the orderingnever changes: ifosfamide most sensitive at all three timepoints, doxorubicin least at all three.

Why any of this is agentic rather than automatic

A pipeline is a fixed sequence. This study was not one, and the record shows where the judgement calls were.

The tumour signature had to be learned from the patient rather than taken from literature, because collagen markers merge osteosarcoma cells with fibroblasts and blood vessels, which are exactly the populations that had to stay separate. The mask over the tissue had to be spatially smoothed, because thresholding a 64-micrometre square at 530 counts follows sampling noise rather than tissue. Cancer cells had to be downsampled to a common depth, because June was sequenced a third as deeply as January and raw coverage would have measured the sequencer. The surface filter had to be inverted from exclusion to inclusion, because otherwise the short list was long non-coding RNAs. The organ-comparison axis had to be dropped, because one target produced a fold change of 314 million by dividing by nearly zero.

None of those were in the plan at the start. Each was found by looking at an intermediate result, noticing it was wrong, and understanding why. That is what the work consists of, and it is why the transcripts are published alongside the tables.

Where this leaves the patient

Ranked first through fourth on eight measurements, then confirmed by a dose calculation: TNC,ITGA10 and B7-H3. Of the three, B7-H3 already has an antibody drug conjugate in phase 2trials in sarcoma, which makes it the one that could be acted on soonest. ITGA10 came out ofthe patient's own data with no prior instruction, and turns out to have a body of preclinical workbehind it arguing for exactly this use.

Ranked 85th of 87: TROP2, the target of two approved drugs, on 0.2% of this person's cancercells.

That gap is the argument for doing this at all. Not that the software found something new,though it did, but that the same eight measurements applied to one person's tissue can tell youwhich of the field's existing drugs would have been aimed at nothing.

What we would do next

More things, in order.

Confirm the top three at protein level. RNA is not protein. Sitting unused in this project's rawdata is a multiplexed proteomics measurement of the same two cores, plus stained images ofthe same sections. That check is available and was not done here.

Re-quantify the protocadherin locus properly. Nine of the 29 discovered targets sit in onetightly packed genomic region where reads are easily misassigned. They score well and we donot believe them.


Which drug would work, and how you would check

The first parts asked what to build a drug against. This one asks a nearer question: of the drugsthat already exist, which would this patient's tumour respond to.

It is also the part where we stop showing you figures we drew and start showing you the tool.Every number below was pulled out of the running application over its own web interface, andevery screenshot is that application in a browser with this patient's data loaded.

The result first

Six chemotherapy and targeted agents, predicted per cell, at three biopsies across ten months.

p11_img012_760x294.png11_drug_response_time.jpg

Lower means more sensitive. Four things fall out of that table.

Ifosfamide is first at all three timepoints. It is the only drug with a negative score in June, and it stays the most sensitive agent through to April. Its score deepens from −0.87 to −1.10.

Pazopanib is consistently among the strongest. It starts marginal, at +0.05 in June, and is clearly negative by January and April. Second place at every timepoint.

Sirolimus tracks just behind it. Third at every timepoint, on the same trajectory.

Doxorubicin and trabectedin are last, and they are last by a wide margin. Both sit above +1.4 in June, and even in April, when every drug has moved down, they are the only two still on the resistant side of zero.

The ordering never changes. Not once, across three samples taken ten months apart and prepared separately.

The obvious objection, and the test for it

We showed that the cell population in these biopsies changes drastically: 21% cancer in June,17% in January, 1% in April, with T cells going from a fifth of the sample to 86%. A scoreaveraged over every cell in a sample that is changing that much may be measuring the changein composition rather than anything about the drug.

So we recomputed the whole thing inside the cancer compartment alone, using the cell-typelabels the application serves alongside the scores.

12_drug_cancer_cells.jpg

The ranking is identical. Same six drugs, same order, whether you average over every cell inthe sample or only over the cancer cells. The ordering is not an artefact of the shifting cell mix.

The absolute level is. Restricting to cancer cells pulls every drug down: cisplatin goes from+0.36 to −0.04 in June, because the immune cells that dominate the sample are scored as moreresistant than the cancer is. So the level of the all-cells score depends on what else is in thetube, and the ordering does not.

The downward drift over time survives the test only partly. Between June and January the driftis about the same size inside the cancer compartment as it is overall, so that much looks real.April cannot be checked, because the application does not serve cell-type labels for that sampleand, it contains too few cancer cells to test anyway.

Now the tool

Everything above came out of the application's own interface. Here is what that looks like.

Three stored sessions, one per data type: the three single-cell biopsies, the two VisiumHDslides, the two Xenium slides. Opening a session shows what has been produced in it.

Twenty-seven artifacts in the single-cell session: eighteen sensitivity maps, three summarytables, three working data files, three others. None of these were made for this write-up.

The request that produced them, and the answer, are still in the thread.

p13_img013_760x452.png

Note what it lists under Outputs: seven spatial heatmaps, seven ds_* viewer layers, and aworking file with the columns saved. Those layers are exactly what the rest of this post switcheson.

Loading a sample draws every cell.

p14_img014_760x452.png

13,045 cells from the January biopsy, coloured by the cell type each was assigned. Now switchthe colour to a drug.

p14_img015_760x452.png

This is the map the section title promised: every one of those 13,045 cells coloured by its ownpredicted ifosfamide sensitivity, live in the browser. Dark is sensitive. Almost the whole map isdark.

The panel at the top right is worth a moment. It reads −1.017, computed over all 13,045 cells bythe viewer at the moment the screenshot was taken. That is the same number as the Januaryrow of the ifosfamide table above, and the same number the agent reported in April. Threeindependent paths to one value.

Switch to doxorubicin, same cells, same colour range.

p15_img016_760x452.png

+1.052, and the map is bright almost everywhere. This is the contrast in the first table, drawnone cell at a time rather than summarised into a number.

What this is and what it is not

These are predictions from a model that reads a cell's expression profile and a molecule'sstructure, and returns a number. It has never seen this patient. It has no access to the clinicalrecord, and we have not compared any of it against what actually happened to this person,because that is not our data to use.

So the honest statement of the result is narrow. Within this model, applied to this patient'scells, the ordering of these six drugs is stable across ten months, across three separatelyprepared samples, and across every cell compartment in the tumour.

A note on limitations

The sensitivity results should be read as a demonstration of what this framework can do, not asa validated clinical prediction. Models of this kind still require broad external validation acrosstumour types, treatment settings and independent patient cohorts, as well as prospectivecomparison with real treatment response. For now, their value here is explanatory: they showhow patient-specific molecular data can be turned into testable drug hypotheses, not whichtreatment a patient should receive.

https://youtu.be/AefQ0Blqcb4

This case study is for research and exploratory purposes only. The results have not been clinically validated and should not be used to guide patient care.

Recommended reads