Grading the Promise, Not the Prose: A Verbatim Ledger for Executive Hedging
Diligence analysts flag management hedging by feel on an earnings call; a verbatim, dated ledger of exact quotes, specificity scores, and language drift turns that gut check into evidence you can point to instead of an impression you have to defend.
Grading the Promise, Not the Prose: A Verbatim Ledger for Executive Hedging
Every diligence analyst has the same experience on an earnings call: a chief executive officer who said "we will" last quarter says "we remain optimistic about our ability to" this quarter, and something in your gut flags it as a downgrade in confidence before a single number has changed. The trouble is that flag lives in your gut, not in a file. It doesn't get dated, it doesn't get compared against the last four quarters in a consistent way, and it doesn't survive analyst turnover on a deal team.
We've been building the plumbing to turn that gut check into a structured, dated record instead of a one-off note in a call summary — and it's worth walking through what that record actually captures, because it's more specific than "sentiment got worse."
A ledger of quotes, not a summary of tone
The instinct when building something like this is to have a large language model — an AI system trained to read and generate text — read a transcript and hand back a paragraph: "management's tone was more cautious this quarter." That's not useful for diligence work, because it can't be checked and it can't be compared apples-to-apples across companies or time.
What we're building instead is a ledger where each row is one verbatim quote, not a summary. For every claim a management team makes about the business, we're capturing:
- The exact words, plus the sentence before and after it, so the quote can't be read out of context.
- A specificity score — did the executive commit to a number and a date, or to a feeling?
- Whether a superlative was used — "industry-leading," "best-in-class," and similar language get flagged as their own signal, since superlative density tends to rise exactly when specifics fall.
- Drift from the prior statement — how this quarter's version of a given claim compares to what was said about the same initiative last quarter.
- A language-delta score — a more granular measure of how the wording itself shifted, independent of whether the underlying commitment technically stayed "on track."
None of that requires a person to remember what was said four quarters ago. The point of keeping it verbatim and append-only — new quarters add rows, nothing gets overwritten or rewritten in hindsight — is that you can go back and audit exactly what was said, when, and how the language moved before you draw a conclusion from it.
Initiatives get a lifecycle, not a mention count
Separately, and feeding into the same picture, named initiatives — a product launch, a cost-cutting program, a stated margin target — are tracked through an explicit lifecycle: announced, then delivered, abandoned, or delayed. That's a meaningfully different question than "did they mention it in the shareholder letter," because initiatives quietly stop being mentioned all the time without ever being marked as abandoned. Tracking the lifecycle state directly, rather than inferring it from silence, is what lets a "quietly dropped initiative" show up as a data point instead of something you'd only catch by re-reading last year's letter side by side with this year's.
Roll that lifecycle data up across every initiative a company has ever announced, and you get a separate, cumulative figure — a lifetime delivery rate — that answers a different question than any single quarter can: across everything this management team has ever committed to publicly, what fraction actually got delivered? That's a batting average, not a snapshot, and it's the kind of number that's tedious to reconstruct by hand across a multi-year coverage history but straightforward to maintain once every promise is logged as it's made.
What the ledger is built to catch
The quote-level data — specificity, superlative use, drift, language delta — is designed to feed four named patterns we look for directly: promise decay (a commitment getting vaguer over successive quarters without being retracted), abandoned initiatives (dropped without acknowledgment), moat laundering (reframing an eroding competitive advantage as a deliberate strategic choice), and acquisition deflection (attributing an organic miss to portfolio changes). Naming the pattern is the easy part; the reason to build the verbatim ledger first is that none of those four patterns can be scored credibly from a paragraph of vibes — they need a dated trail of exact language to point to as evidence.
Where this actually stands today
In the interest of not overselling a work in progress: the verbatim quote-and-drift ledger described above is the newer half of our scoring pipeline, and it's currently in calibration against a small set of reference companies rather than running across full coverage. Today's live reports already track claims and produce dimension-level scores, just through an earlier, more monolithic version of the pipeline. The granular, quote-by-quote ledger is the piece we're rolling out next — deliberately slower, because a promise-tracking record is only as trustworthy as its worst false positive, and we'd rather calibrate it against a handful of companies until the variance is tight than ship it broadly and have it flag hedging that isn't there.
For a diligence workflow, the practical takeaway is this: the difference between "management sounded less confident" and "management's stated timeline for initiative X drifted from a committed date last quarter to a qualified 'working toward' this quarter, and here's the exact sentence" is the difference between an impression you have to defend in an investment committee meeting and a citation you can just point to.
