A bulk RNA-seq profile from an acute leukemia sample combines expression from malignant and non-malignant cells. It therefore reflects both the sample's cell-state composition and cell-intrinsic gene regulation that cannot be explained by composition alone. We address two questions:
Question 1. Which normal haematopoietic differentiation states does each acute-leukaemia sample resemble?
Question 2. Is the transcriptomic signal of each molecular subtype primarily explained by its differentiation-state composition, or does expression retain subtype information after composition is accounted for?
We use the remaining expression signal as an operational measure of expression beyond composition. This distinction is based on held-out prediction and is interpreted at the subtype level.
The analysis needs two things: a set of bulk tumour transcriptomes to explain, and a reference describing what normal blood-forming cells look like. This section assembles both.
The bulk samples. We collected 1,714 transcriptomes spanning the three acute leukemias: 1,236 AML profiles from the AMLmap multi-cohort study,1 328 B-ALL and T-ALL profiles from GSE227832,2 and 150 diagnosis T-ALL profiles, one per subject, from St Jude Cloud feature-count files.3 All samples were placed in a common gene-symbol matrix.
The normal-cell reference. Deconvolution can only recognise the cell states supplied to it, so the reference bounds which states the analysis can detect. We built ours around BoneMarrowMap,4 because it resolves the marrow hierarchy at the granularity this analysis needs and was itself constructed for AML differentiation-state mapping. It supplies 32 marrow states — stem and progenitor, myeloid, erythroid, megakaryocytic, B-cell, and mature T/NK — but, being a marrow atlas, it stops before the thymus, where T-ALL arises. To it we added the three earliest stages of intrathymic T-cell development — ETP-T, DN-T and DP-T — from the Park human thymus single-cell atlas,5 giving one 35-state reference. Without them, no thymic state is available to a T-ALL sample, and its weight is assigned instead to the marrow states that resemble it most closely.
Plasma reference state. This supports one visual coordinate system, but does not establish
perfect separation of all 35 signatures in bulk mixtures; Part 2's AML negative control tests that
harder question.We compared NNLS, MuSiC and InstaPrism using the same bulk matrix, reference cells, gene filters and 35-state labels. The methods produced visibly different composition estimates, so the downstream method could not be selected from visual plausibility alone.
Sixteen Leucegene AML specimens have both bulk RNA-seq and single-cell RNA-seq from the Leucegene surfaceome study (GSE241989).10 We assigned each single cell to the same reference states and calculated a single-cell composition for each specimen. We then compared that composition with the estimate obtained from its matched bulk sample.
The primary evaluation used five broad non-thymic lineages. A second evaluation used 32 individual non-thymic states. Thymic states were excluded because these matched specimens are AML and therefore provide a negative control for thymic assignment.
The negative control does not pass. An AML sample should place no weight on ETP-T, DN-T or DP-T. Across the 1,058 AML samples with an annotated subtype, InstaPrism puts a mean of 16.0% of their composition on those three states (median 15.6%). This is bleed between the two atlases that make up the reference — the thymic profiles are close enough to cycling and peripheral marrow programs to absorb weight from them — and not a property of the tumours. It is the reason the benchmark below is scored on non-thymic states only, the reason the map's second panel z-scores each state across subtypes, and the reason we read the T-ALL thymic enrichment as a contrast between rows rather than as an absolute fraction in any one sample. All three methods assign thymic weight to these samples; InstaPrism assigns more than NNLS or MuSiC while being more accurate on the states that can be present.
A mean correlation does not distinguish a method that is accurate on every specimen from one that is accurate only on average, and it does not characterise the errors. Correlation is also insensitive to scale: a method that systematically pulls every fraction toward the middle, never calling a lineage dominant or absent, can still correlate well with the reference. To separate these failure modes we plot each specimen separately and fit a line through its five lineages. The correlation measures whether the method ranks the lineages correctly; the slope of that line measures whether it reproduces their magnitudes. The two measures can disagree.
We selected InstaPrism for the remaining analyses because it had the highest mean accuracy at both resolutions and a statistically supported advantage at broad-lineage resolution. We do not interpret the fine-state advantage over MuSiC as conclusive.
We applied InstaPrism to all 1,714 bulk samples and summarized the estimated state composition for each molecular subtype. Absolute fractions describe the dominant reference states. State-wise standardized values describe which states are relatively enriched or depleted for a subtype compared with the other subtypes.
Panel a reports absolute, within-row normalized resemblance weights, and therefore locates each subtype on the normal reference map. Panel b reports the same weights after z-scoring each column across the displayed rows, which removes the background shared by all subtypes and isolates relative enrichment or depletion. A strong signal in panel b can consequently arise from a modest absolute weight, because the z-score is a contrast between subtypes rather than a cell fraction or an effect size. Thick black rules separate AML, B-ALL and T-ALL, and thinner white rules separate molecular subtype classes within each disease.
AML. The PML-RARA row is concentrated in the myeloid/GMP–neutrophil region of the reference, which is compatible with the defining promyelocytic differentiation arrest of acute promyelocytic leukemia.11 The two core-binding-factor subtypes, RUNX1-RUNX1T1 and CBFB-MYH11, also occupy the primitive-to-myeloid continuum, but with visibly different within-lineage profiles. That is biologically plausible: both fusions disrupt a complex required for myeloid differentiation, yet they have distinct cooperating-mutation spectra and clinical phenotypes.12 This agreement is a useful compatibility check; it is not an independent validation of the inferred proportions.
B-ALL. The B-ALL subtypes concentrate around the Pro-B and Pre-B columns. This is consistent with the developmental nature of B-cell precursor ALL. In particular, single-cell profiling of ETV6-RUNX1 leukemia found that its transcriptional state most closely resembles pro-B cells, with an arrested B-lineage program.13 The map adds a common coordinate system in which this pattern can be compared directly with other B-ALL subtypes and with AML, while retaining the important caveat that a normal-state match is not a literal developmental timestamp.
T-ALL. T-ALL rows are enriched for ETP, DN and DP thymic programs rather than marrow B-cell or myeloid states. This agrees with the established view that T-ALL molecular subgroups reflect arrest at different stages of thymocyte development; for example, TLX-class and TAL/LMO-class cases have distinct cortical-stage associations.14 Because the thymic cells come from a separate reference dataset, we interpret the relative T-ALL enrichment and the ordering of T-ALL rows more strongly than any absolute thymic weight in an individual sample.
The T-ALL contrast reproduced in the two independent cohorts. The median combined ETP-T, DN-T and DP-T weight was 0.79 in the 19 GSE227832 T-ALL samples and 0.63 in the 150 St Jude T-ALL samples, compared with 0.15 in AML and 0.32 in B-ALL. Thymic weight alone separated T-ALL from AML with AUC 0.992 (95% bootstrap interval 0.983–0.998) and from B-ALL with AUC 0.935 (0.908–0.960). This replication shows that the relative thymic signal is not confined to one T-ALL cohort. It does not remove the non-zero baseline in the other diseases, so it strengthens the between-group contrast rather than validating any absolute fraction.
The new result is therefore not the rediscovery that PML-RARA is promyelocytic or that T-ALL is thymic. It is the subtype-level comparison across AML, B-ALL and T-ALL on one normal differentiation axis. Panel B makes the finer distinctions explicit—for example, differences among AML entities, myelodysplasia-related mutations and cytogenetic subtypes that would be obscured by their shared primitive/myeloid background. These finer associations are hypothesis-generating until tested with genotype-resolved single-cell or orthogonal immunophenotyping data.
Individual samples vary considerably around the subtype means. The fusion subtypes are the tightest: the summed myeloid fraction of PML-RARA samples is 0.53 with a standard deviation of 0.11, and CBFB-MYH11 is 0.54 ± 0.14. Mutation-defined subtypes are broader — NPM1 is 0.46 ± 0.22 and the myelodysplasia-related mutations SRSF2 and ASXL1 are 0.41 ± 0.22 and 0.43 ± 0.20 — with individual samples ranging from almost entirely primitive to almost entirely monocytic. Selecting a single subtype in the map above shows this directly: the subtype means summarise a distribution rather than a stereotyped differentiation state.
For each subtype, we compared two held-out models:
Composition only: predict whether a sample is subtype-positive (1) or subtype-negative (0) using 35 estimated differentiation-state fractions.
Composition + genes: predict the same 0/1 outcome using the 35 fractions plus 15 subtype-specific genes.
The difference between the two models quantifies how much subtype detection improves when expression is added beyond composition. We report that improvement as the expression share: 0% means composition accounts for nearly all above-chance detection, and 100% means genes account for nearly all of it.
Each subtype is tested separately. Samples positive for the subtype are coded 1; other samples from the same disease group are coded 0. The outcome is binary, while the classifiers use multiple predictors.
Held-out samples never influence filtering, expression cutoffs or gene selection. The 15-gene cap is a common measurement budget, not a curated signature. AML uses leave-one-cohort-out evaluation; ALL uses five-fold cross-validation.
AUC is a ranking score: 0.50 is chance, while 0.80 means a random subtype-positive sample is ranked above a random subtype-negative sample 80% of the time.
For SF3B1, imagine 1,000 comparisons: chance gives 500 correct rankings, composition gives 605, and
composition plus genes gives 788. Composition adds 105 rankings above chance and genes add another 183,
so the expression share is 183 / (105 + 183) = 64%.
Most of SF3B1's above-chance detection appears after genes are added. For PML-RARA, the expression share is only 5% because composition already identifies the subtype well.
A high expression share can occur when composition starts near chance, even if the absolute gene gain is modest. Across interpretable subtypes, the share correlates −0.74 with composition AUC and +0.85 with the absolute gain from genes. The figure above therefore shows both model scores alongside the ratio.
Adding clinical blast percentage to the AML composition model made little difference across 22 subtypes: the mean absolute AUC change was 0.006 for composition and 0.004 for the joint model. For SF3B1, the AUCs changed from 0.605 to 0.617 and from 0.788 to 0.786; for STAG2, from 0.615 to 0.611 and from 0.847 to 0.836. Blast burden therefore did not explain the expression gain.
RUNX1-RUNX1T1, PML-RARA and CEBPA were predominantly detected from composition. SF3B1, STAG2 and U2AF1 showed larger expression contributions. This pattern was not uniform across mutation classes: SRSF2, a splicing factor, sits at 30% and therefore nearer the compositional end than several cytogenetic subtypes. Splicing- and chromatin-associated subtypes are not, as a class, expression-dominated.
The clearest structure in the panel is not at the level of individual subtypes but at the level of the subtype classes. Across the 27 subtypes with an interpretable share, the fusion- and aneuploidy-defined entities of AML and B-ALL sit low (median share 19% and 18%), while the myelodysplasia-related point mutations sit high (median 54%). The two groups differ mainly in where composition starts: the median composition-only AUC is 0.889 for AML entities and 0.854 for B-ALL subtypes, against 0.615 for the MDR mutations. A translocation that arrests the blast population at a particular point of the hierarchy is therefore detectable from the fractions alone. A splicing or cohesin mutation does not appear to fix a differentiation stage in the same way, and the fractions leave most of the subtype undetected.
The added-expression signal also survives a cross-cohort direction test. Within AML, we regressed expression on differentiation-state composition and cohort, selected the 50 strongest residual genes using all other eligible cohorts, and then asked whether their effect directions agreed in the held-out cohort. Sign agreement was 98% for STAG2, 93% for U2AF1 and 85% for SF3B1; only 0–2% of their top residual genes were technical features or top reference-state markers. This makes a simple cohort or residual cell-state explanation less likely. Because the adjustment still relies on estimated composition, it supports—but does not prove—a within-state transcriptional effect.
Four subtypes fall off the axis entirely, because adding genes made held-out prediction worse: del(12p), del(13q), i(17q)/del(17p) and T-ALL NOS. In each, composition alone already predicted at 0.56–0.82, and the 15 selected genes did not generalise across held-out cohorts or folds. We read these as subtypes with no expression signal that survives a 15-gene budget and a cohort shift, rather than as evidence that composition is a complete description of them. They are hidden by default in the figure and excluded from every summary statistic above.
A high share also does not imply that a subtype is well detected. EZH2 has a 52% share, but its joint AUC is 0.561: genes supply most of what little separation exists, and there is very little. The same applies to BCOR (joint AUC 0.647, share 10%). The share indicates where a subtype's signal originates, and the two AUCs indicate whether that signal is large enough to interpret; panel A reports both. Among the mutations, only SRSF2 (0.939) and STAG2 (0.847) combine a substantial expression contribution with a joint model that classifies well.
Each fraction is a ratio of two estimated AUCs, and neither is measured without error. We therefore interpret a subtype's position along the axis rather than the particular percentage attached to it. A difference of a few percentage points between adjacent subtypes does not support a biological distinction. The ordering is at least not an artefact of subtype prevalence: the share is uncorrelated with the number of carriers (r = 0.02, n = 27), so the mutations do not sit high merely because they are rarer than the fusions.
Acute-leukaemia subtypes differ not only in how strongly they alter the transcriptome, but in what kind of transcriptomic signal they leave. By placing 1,714 AML, B-ALL and T-ALL samples on the same 35-state haematopoietic reference, we recovered a coherent developmental map: AML subtypes distributed along the primitive-to-myeloid continuum, B-ALL around Pro-B and Pre-B states, and T-ALL around early thymic states. The familiar locations of entities such as PML-RARA and ETV6-RUNX1 provide biological checks on the map. The contribution is the shared coordinate system itself, which makes differences among diseases and among molecular subtypes directly comparable rather than treating each leukaemia as a separate landscape.
That common map revealed a systematic division in where subtype information resides. Across the 27 subtypes with interpretable estimates, fusion- and aneuploidy-defined entities were identified largely by differentiation-state composition: the median expression-beyond-composition share was 19% for AML entities and 18% for B-ALL subtypes. PML-RARA, RUNX1-RUNX1T1 and CBFB-MYH11, for example, were already classified with held-out AUCs of 0.93–0.96 from the 35 state weights alone. In contrast, myelodysplasia-related point mutations had a median expression share of 54%. STAG2, SF3B1 and U2AF1 began close to chance from composition and gained most of their detectable signal when genes were added. The same contrast was visible within B-ALL, from composition-led ETV6-RUNX1 (11%) to expression-led iAMP21 (66%). Thus, the main finding is an ordered spectrum: some lesions are recorded primarily in where leukaemic cells sit in the differentiation hierarchy, whereas others are recorded primarily in what they express beyond that position.
Two checks strengthen that interpretation. The early-thymic contrast reproduced in two independent T-ALL cohorts and separated T-ALL from AML and B-ALL with AUCs of 0.992 and 0.935, respectively. The residual expression directions of SF3B1, STAG2 and U2AF1 also agreed across held-out AML cohorts after adjustment for composition and cohort. Together, these checks show that the principal signals are not confined to a single cohort or explained by cohort structure alone.
This spectrum is not a claim that every fusion is compositional or every splicing or chromatin lesion is expression-dominated; SRSF2 is an important counterexample. Nor are the reported shares literal biological fractions. Their defensible interpretation is comparative: the class-level separation and the broad ordering of subtypes are more informative than small differences between adjacent percentages. Within that boundary, the analysis identifies a concrete biological distinction that bulk expression alone normally obscures and supplies testable hypotheses for genotype-resolved single-cell studies: lesions at the compositional end should chiefly redistribute malignant cells across developmental states, whereas lesions at the expression end should remain distinguishable among cells occupying the same state.
The thymic negative control does not pass. The 16 matched specimens are AML, so any weight they place on the three thymic states is error. On the merged 35-state reference, InstaPrism assigns AML samples a mean of 16.0% of their composition to ETP-T, DN-T and DP-T (median 15.6%), which is biologically impossible in marrow or blood. This is bleed between the two atlases, not a property of the tumours. It is why the benchmark in Part 2 is scored on non-thymic states, why panel b of the map z-scores each state across subtypes, and why we read T-ALL's thymic enrichment as a contrast against the other rows rather than as an absolute fraction. It does not propagate into Part 4, where every subtype is compared against same-disease controls that carry the same bleed — but the absolute thymic weight of any individual AML sample is not interpretable. The replicated T-ALL contrast shows that the signal rises well above this baseline; it does not turn the weights into calibrated cell fractions.
The matched benchmark selects a method; it does not validate the compositions. Two limits compound. The single-cell "truth" is itself a projection of each cell onto the same 35-state reference using the same marker genes as the bulk deconvolution, so the two sides share their errors: this supports the ranking of NNLS, MuSiC and InstaPrism, but inflates the absolute correlations. And the single-cell specimens come from a surfaceome study centred on primitive populations; we have not confirmed from its protocol that the libraries were whole viable cells rather than enriched, and an enriched input would bias the census we score against. The benchmark also contains only AML, and therefore says nothing directly about B-ALL or T-ALL composition accuracy.
The expression share is a ratio, and its denominator depends on composition performance. Because composition's above-chance signal sits in the denominator, a subtype that composition cannot detect will receive a high expression share even when expression adds a modest absolute gain. Across the interpretable subtypes the expression share correlates −0.74 with the composition AUC, and the subtypes at the expression-dominated end are also the least detectable overall (total AUC 0.77–0.85, against 0.95–1.00 for the fusions). The share is not solely an artefact of the denominator, since it also tracks the absolute expression gain (r = +0.85); panel B prints both AUCs, which separates a large expression contribution from a failure of composition. Estimates near 0 or 1, and those resting on few carriers, are correspondingly less stable, and no uncertainty interval is attached to any individual fraction.
Finally, the deconvolution output measures similarity to normal reference states. It is not tumour purity, not a histological cell fraction, and not a developmental timestamp.
The bulk discovery set draws on two studies. AMLmap contributes five AML cohorts, and the Nordic ALLIUM study contributes the molecularly annotated ALL RNA-seq subset deposited as GSE227832. The normal-cell reference combines two single-cell atlases: we retained 32 marrow states from BoneMarrowMap and appended three early thymic states from the Park thymus atlas. For the matched evaluation in Part 2, sixteen Leucegene AML bulk samples were paired by patient identifier to single-cell profiles from GSE241989.
The two atlases were merged at the level of cells, not of published signatures. From the Park thymus
atlas we kept only the pre-selection stages that BoneMarrowMap lacks and mapped its labels onto three
states: DN(early) became ETP-T, DN(P) and DN(Q) were merged into
DN-T, and DP(P) and DP(Q) were merged into DP-T. The (P) and
(Q) suffixes denote proliferating and quiescent cells at the same developmental stage, so
merging them removes a cell-cycle distinction rather than a developmental one. We excluded
αβT(entry), a broad transitional state that otherwise absorbs weight from its
neighbours, and we did not import thymic single-positive CD4 or CD8 cells, because BoneMarrowMap already
contributes mature T and NK states and two near-identical sets of columns would let the fit divide weight
between them arbitrarily.
For the cell-level visualization, mature CD4, CD8 and regulatory T cells from both atlases were retained temporarily as matched bridge populations. We calculated a 50-dimensional expression space and aligned the mean and shrinkage-covariance structure of the Park bridge cells to their BoneMarrowMap counterparts. Park bridge cells were then removed before the UMAP was fitted and displayed. Cross-atlas neighbour mixing among the three bridge groups and within-state neighbour purity among the displayed cells were calculated as integration diagnostics. This correction is used only for the reference visualization; the deconvolution uses the uncorrected state-specific signatures described next.
Every state was then subsampled to 150 cells and summarised as a counts-per-10,000 signature. Ribosomal,
mitochondrial and sex-linked genes were removed; cell-cycle genes were deliberately retained, since
proliferation is part of what distinguishes progenitor states. InstaPrism and NNLS use all genes shared with
the bulk matrix; MuSiC uses a panel of the 20 most specific genes per state.reference_merged_fine.npz; built by export_merged_native.py.
For each bulk sample, we fit its linear counts-per-10,000 expression profile to a matrix of state-specific reference profiles, using the union of the 200 most specific genes per state. We down-weighted highly expressed genes, constrained all coefficients to be non-negative, and rescaled the fitted coefficients to sum to one. These are therefore normal-state resemblance weights, rather than directly observed cell fractions. NNLS is the standard constrained least-squares estimator;6 MuSiC uses multi-subject single-cell weighting;7 and InstaPrism is a fast fixed-point implementation of the BayesPrism model.89
For the pan-lineage stress test, we summed ETP-T, DN-T and DP-T weights in each sample and bootstrapped the AUC for T-ALL versus AML and T-ALL versus B-ALL. The two T-ALL cohorts were summarized separately.
For the residual-expression audit, log-CPM expression was adjusted within disease for 12 principal components of the estimated composition and for cohort indicators. For each AML subtype, the 50 largest residual effects were selected without the cohort being evaluated; replication was the proportion retaining the same direction in that held-out cohort. Ribosomal, mitochondrial, sex-linked and cell-cycle genes were flagged as technical controls, and the top 20 specific genes for each reference state were flagged as state-marker controls.
Pipeline in acute_leukemia_map/writeup_pipeline/. The standardised merged-35
analysis is driven by a shared config module (merged35.py). The matched-method benchmark is
generated by story_completion_stats.py; that script also produces the pan-lineage stress test.
The subtype composition/expression model is generated by
comp_vs_intrinsic_merged.py, and story_completion_expression.py produces the
cross-cohort residual audit. The integrated cell-level reference figure, exported embedding and QC table
are produced by acute_leukemia_map/integrate_reference_atlas.py. Method estimates:
instaprism_theta_merged.csv, music_theta_merged.csv,
nnls_theta_merged.csv.
Every figure in this article is drawn in the browser from the same inputs as the corresponding
matplotlib figure in the pipeline, which remains the archival version.
export_figure_data.py writes
data/matched.js (the 16-specimen benchmark) and data/intrinsic.js (every subtype
in comp_vs_intrinsic_merged.json); export_persample_composition.py writes
data/persample.js. Drawing is shared through lib/figlib.js, which mirrors the
palette and colormaps of src/natfig.py. In Section 4, subtypes whose composition AUC does not
exceed chance, or for which expression did not improve held-out AUC, are marked † and drawn faded: the
intrinsic fraction printed for them is a clamp, not an estimate.
Every figure carries a Data line beneath its caption offering the numbers it draws as CSV: the 35
reference states and their lineages; the benchmark's method summary, per-specimen correlations and
paired-bootstrap contrasts; the 16 matched specimens' single-cell and predicted compositions at both
resolutions; the 35-state composition of all 1,714 samples together with subtype membership; and the two
AUCs, expression share and mean lineage composition of every subtype. The files are assembled in the
browser from the same arrays the canvases read, so a download cannot disagree with the panel above it.
Compositions are written as fractions summing to one, and the † caveat applies to the
expression_share column exactly as it does to the figure — the reason is carried in that
file's reliability_note column.