Drug-Level Sum-of-the-Parts · rNPV

Pharma Valuation Pipeline

A point-in-time engine that values a pharmaceutical company one drug at a time — each marketed drug as a lifecycle DCF, each clinical asset as a risk-adjusted NPV, plus an R&D renewal engine for the drugs not yet born — then bridges to a per-share intrinsic value.

HOW TO READ THIS Every stage separates the live production path — what a valuation actually runs on today — from the feature-tuned models calibrated alongside it, which run in parallel and switch on once they prove out. Empirical results and backtests live in the dashboard.
Live production path Calibrated model (built, gated off) Deterministic LLM agent Blocking gate
00The full pipeline, at a glance

One company, one as-of date, eleven stages. Data is acquired point-in-time, a three-agent LLM trio extracts the drug portfolio under deterministic gates, the deterministic engine values three legs, and a QC layer quarantines anything it cannot defend. Hover-free, scroll down for the detail behind any stage.

① POINT-IN-TIME DATA ACQUISITION every input dated as of the valuation date 1 · Classify pharma gate + archetype 2 · Data snapshot WRDS · Refinitiv · SEC · CT.gov · yfin 3 · Web research 6 parallel · non-blocking 4 · Financials WRDS fundamentals + SEC filings 5 · WACC CAPM + capital-risk overlay Reproducible run record every run is re-verifiable end to end ② LLM EXTRACTION TRIO · validate → retry → reconcile 6 · Drug-Portfolio Agent marketed drugs · revenue · LOE 7 · Pipeline Agent clinical assets · phase · peak est. 8 · Reviewer (conditional) accepts only if it strictly reduces errors Reconciliation gates vs SEC 10-K Product Sales · vs CT.gov ③ DETERMINISTIC VALUATION ENGINE — three independent legs Marketed DCF per drug: growth→peak→plateau→erosion gated Gordon tail · Medicare-slice IRA ∑ marketed EV Pipeline rNPV PoS-weighted revenue − staged cost live: literature PoS  ·  calib: ML model ∑ pipeline EV Platform / renewal value of drugs not yet born credits only R&D that pays off platform value Equity bridge EV − overhead − net debt (special handling if distressed) ⇒ $ / share ④ REPORT & QC Investment memo with a full audit trail · QC flags any number it cannot defend rather than deleting it · each company carries a data-quality score. This page covers how a value is built; how well it performs is measured separately in the dashboard.
The live path runs end-to-end deterministically; LLM agents touch only extraction (stages 6–8). Feature-tuned calibration models sit alongside the engine and are detailed in each section below.

Three-leg Sum-of-the-Parts

Enterprise value = Σ marketed-drug DCFs + Σ pipeline rNPVs + an R&D renewal-engine platform leg, minus corporate overhead. No portfolio-level revenue assumption is ever needed — value is built bottom-up, drug by drug.

Live vs. calibrated, everywhere

The production engine is deliberately conservative: it runs on well-established literature parameters. A parallel set of feature-tuned ML models is fitted and benchmarked, but each is gated OFF until it clears documented promotion criteria.

Honest by construction

Strict point-in-time data discipline, analyst consensus used only as a reference (never an input), and a quality-control layer that flags anything it cannot defend rather than quietly shipping it.

01Point-in-time data & discount rate

Every input is anchored to the valuation date — the model only ever sees what was knowable then. A historical valuation reads from a frozen snapshot of that date's data, never from today's prices or news, so a backtest can never accidentally "see the future" and a live run can never quietly drift.

WRDS primary

Compustat fundamentals, CRSP prices/returns, IBES point-in-time consensus. The reproducible backbone, with as-of dates that avoid look-ahead.

Refinitiv

Current market data and analyst consensus for the live valuation. Consensus is used only as a reference point to compare against — never as an input the model fits to.

SEC EDGAR

10-K / 10-Q filings, the drug-level Product Sales table (the source of truth we reconcile drug revenue against), and patent-expiry & risk disclosures. Machine-readable filing data fills any gaps in WRDS.

ClinicalTrials.gov · market data

ClinicalTrials.gov confirms which pipeline programs existed and their phase; market feeds supply price, shares and beta. Companies are matched to filings even after a ticker change or acquisition via a hand-verified identity map.

The revenue evidence table is segment-level — and says so.

After Visible Alpha's drug-level feed was retired (credentials lapsed mid-2025), the assembled revenue evidence is built from WRDS + Refinitiv segments and is explicitly tagged is_drug_level = false. Drug-level revenue comes from filings and the agents, reconciled against SEC Product Sales — not from a vendor drug feed.

The discount rate (WACC) is built in two stages deterministic

1 · Textbook cost of capital

cost of equity = risk-free + β · equity-risk-premium
WACC = equity-weight·(cost of equity) + debt-weight·(after-tax cost of debt)

Standard CAPM, with the company's betaHow much the stock moves with the market. It's shrunk toward the market average to damp noisy single-stock estimates. stabilized and the cost of debt inferred from credit quality when the reported figure looks unreliable.

2 · Pharma risk overlay

A bounded add-on for company-specific fragility — small size / thin liquidity, high volatility, short cash runway, and a stretched balance sheet. It can only move the rate within tight limits, so no single factor dominates.

This overlay-adjusted rate is what discounts every leg of the valuation.

02Marketed-drug DCF

Each approved drug is its own discounted cash flow across a four-phase lifecycle. The explicit horizon is adaptive — (LOE − start) + 12 years, capped at 40 — and a continuation tail is added only where it is economically justified.

Revenue lifecycle — per drug years since launch growth peak · plateau IRA cut 40% × Medicare slice (~36%) LOE erosion — keeps decaying toward a small durable floor (no flat shelf) gated Gordon tail Gordon tail is conditional Added only for durable curves (biologic / orphan / gene therapy), g ≤ 0, cap = max(WACC, 6%). Sharp small molecules, complex generics & cyclical drugs (e.g. COVID) get no tail. After-tax cash flow revenue × profit margin × (1 − tax), times the company's share, discounted to today.
Eight modality-specific erosion curves drive the post-LOE decline; the curve continues geometrically toward a durable floor rather than holding a flat shelf.

Profit margin — live production

what a valuation runs on today

A single 70% after-tax margin on branded drug revenue (roughly 20% cost of goods + 10% selling & admin), applied uniformly.

  • Simple and robust — no risk of mis-labelling one drug's economics
  • Lower margins are still applied to non-drug lines (diagnostics, devices, plasma…)

Per-drug margin — calibrated in parallel

a finer margin per drug type

A margin table specific to each therapeutic-area × drug-type combination (~30 in all) — e.g. an oncology biologic carries different economics than a cardiovascular pill.

  • More precise once the drug-type classifier is fully audited
  • Runs alongside; the production number stays on the flat margin for now

Price negotiation is partial, not full

The Inflation Reduction Act lets Medicare negotiate prices, but only Medicare-covered sales are affected. The model applies the cut to just the ~36% of revenue that is Medicare-exposed — not the whole drug — for the years a drug is eligible.

Automatic erosion curve

If no erosion profile is supplied, one is chosen from the drug type: pills erode fast, biologics follow a slower biosimilar curve, complex cell/gene therapies slower still, and generics fastest of all.

Horizon fits the drug

Each drug is projected explicitly until ~12 years past patent expiry (capped at 40) — long enough to capture durable franchises without assuming a dying pill lasts forever.

03Pipeline value & probability of success

Most clinical drugs never reach market, so a pipeline asset is valued by its risk-adjusted NPVRisk-adjusted net present value: the drug's future cash flows weighted by the chance it actually gets approved, minus the cost still needed to develop it. — future revenue weighted by the chance of approval, minus the development cost still to come. The chance of approval is where the live path and the calibrated model deliberately differ.

value = present value of ( approval probability × revenue × margin × tax )  −  present value of ( remaining development cost )

Two design choices keep it honest. Abandonment floor: a program is never carried at negative value — if the economics turn negative it is simply dropped, exactly as a developer would. Staged cost: the current phase costs 100¢ on the dollar, but later phases are weighted by the chance of ever reaching them — so you only "spend" Phase III money if you get to Phase III.

Live PoS — literature base rates production

what every shipped valuation uses today

A curated lookup of phase-transition probabilities across 17 therapeutic areas (Wong 2019 + BIO 2022 + Nature 2025), combined with eight additive modifiers applied at the relevant gate.

  • Well-established, externally citable, stable
  • Modifiers: selection biomarker +15pp · breakthrough +10pp · strong prior data +12pp · orphan / fast-track / first-in-class / antibody +5pp · failed prior Ph3 −20pp
  • Cumulative PoS clamped to [1%, 99%]

Calibrated PoS — feature-tuned model in parallel

a model trained specifically for this

Rather than read a published average, this estimates each gate (Phase 1→2, 2→3, 3→Approval) from the drug's own features — learned from 50,749 historical trial transitions (FDA + ClinicalTrials.gov, 2000–2025).

  • Features include therapeutic area, modality, trial design & size, sponsor track record, and biomarker / orphan status
  • Captures, e.g., that a biomarker-selected oncology trial is a very different bet from an unselected one
  • Runs alongside the live path; switches on once it clearly beats the published rates
Two estimates of the same number

The literature tables give a robust, well-cited average for a drug's category; the feature-tuned model gives a drug-specific estimate from its trial design and sponsor history. Today the production engine values pipelines on the published rates and runs the calibrated model alongside for comparison — a deliberately conservative default that upgrades to the learned model when it has earned it.

Live phase-transition rates — by therapeutic area PRODUCTION PATH Probability of advancing each gate; cumulative = chance of reaching approval from Phase I. The feature-tuned model re-estimates these same cells from drug-level features. Therapeutic area P1→2 P2→3 P3→NDA NDA→Appr Ph1 → approval note Oncology — solid tumor 57.6%32.7%60%90% 10.2% heterogeneous tumors Oncology — hematologic 65%45%62%92% 16.7% better-characterized Immunology / autoimmune 69.8%45.7%63.7%91% 18.5% rheum · derm · GI Cardiovascular 73.3%65.7%62.2%90% 26.9% large outcomes trials Rare disease (non-onc) 78%55%65%95% 26.5% orphan pathways Infectious disease 70.1%58.3%75.3%92% 28.3% clear endpoints Vaccines 76.8%58.2%85.4%93% 35.5% highest overall LoA Gene therapy (orphan) 80%60%60%95% 27.4% durable, slow erosion 17 therapeutic areas total · sources: Wong 2019, BIO / Informa 2011–2020, Nature Communications 2025, Tufts NEWDIGS 2023 · the calibrated model re-estimates these same cells from each drug's own features.
live the literature table is the shipped PoS source  ·  calibrated the feature-tuned model produces a parallel, drug-specific estimate for the audit trail until promotion.

Phase II is the killer

Across most areas the Phase II→III transition has the lowest pass rate (≈30–50%) — small-trial efficacy signals that fail to replicate. The model makes Phase II assets worth dramatically less than Phase III.

Biomarker effect

Biomarker-selected programs approve ~3× as often (25.9% vs 8.4%, BIO 2022). The live path encodes this as a +15pp modifier; the calibrated model learns it directly from the trial-design feature.

Staged cost-to-complete

Development cost varies by area (CNS ~$750M, oncology ~$650M, rare disease ~$300M from Phase II) and is probability-weighted across remaining phases — not charged up front.

04Platform — the R&D renewal engine

A pharma company is a renewal engine: it keeps spending R&D to launch new drugs that replace the ones going off patent. This leg captures the value of those not-yet-existing drugs — and, crucially, awards nothing to a company whose research has historically destroyed value.

The core idea: only reward R&D that earns more than it costs

Each year of R&D is credited only with the economic profit it generates — the amount by which a dollar of research historically returns more than a dollar. That stream is capitalized into a perpetuity, but only after the named pipeline has matured, so it never double-counts drugs already valued one-by-one. The R&D figure is also taken net of what's already committed to the named pipeline.

Melting ice cube

If a company's R&D has historically returned less than it cost, this leg is floored to zero — a value-destroying research engine earns no credit for "future pipeline," no matter how busy it looks.

No double-counting

Only genuinely new, unborn franchises are valued here. The R&D already spent on named pipeline drugs is excluded, because those drugs are valued individually in the pipeline leg.

Bounded by company type

Capped as a share of the drugs we can already see: large-cap 12%, mid-cap 25%, biotech 70% — halved when the pipeline dominates. Pre-revenue biotechs get none (all their value is in the pipeline leg).

R&D productivity — live production

how much a dollar of R&D returns, by company type

The R&D-return multiple is read from a historical calibration covering 28,762 companies, resolved by how each firm's story actually ended (approval, acquisition, or failure). A present-day valuation uses the most recent year available.

  • Looked up by company type and therapeutic focus
  • Strictly historical — never uses data from after the valuation date

Firm-specific productivity — calibrated in parallel

a model trained per company

Instead of a company-type average, this predicts a specific firm's next five years of approvals and revenue renewal from its own fundamentals — R&D intensity, margins, balance sheet, acquisition spend, and approval track record — learned from 19,919 company-years.

  • Lets a standout R&D engine score above its category average
  • Runs alongside the live path; the production number stays on the company-type calibration for now
Why it's built this way

The hard part of any "future pipeline" estimate is not double-counting drugs you've already valued and not rewarding R&D that doesn't pay off. Crediting only economic profit (returns above cost), starting only after today's pipeline matures, and zeroing out value-destroying research engines handles both — so the platform leg adds genuine going-concern value without inflating the total.

05Equity bridge & quality control

The three legs sum to enterprise value, less a corporate-overhead drag, then standard balance-sheet items bridge to equity and a per-share value. Distressed companies get special handling, and anything the engine cannot stand behind is flagged rather than silently shipped.

Enterprise value → per share MarketedDCFs + Pipeline rNPV + Platform overhead drag = EVenterprise value net debt + preferred+ minority = Equity÷ diluted shares= $/share distressed firmsequity valued as anoption on thebusiness
For a heavily indebted company near insolvency, the equity is treated like an option on the businessWhen debt nearly exceeds asset value, shareholders still hold upside if things recover. Valuing equity as a call option captures that, rather than just flooring it to zero. — capturing the recovery upside shareholders still hold, instead of flooring the value to zero.

Flag, don't delete

The engine checks its own work at every stage. When a result rests on inputs it can't stand behind, it marks the value as not publishable and attaches the caveats — so a weak number is clearly labelled rather than quietly passed off as solid.

One locked configuration

Live valuations, historical replays, and re-runs all use one identical, version-locked set of parameters — so a result can't drift just because it was produced on a different day or path.

Data-quality score

Every company carries a transparency score for how much of its revenue is pinned to specific, identified drugs versus thin or inferred data — computed only from what's observable, never from the stock price.