Boom Leverage

This is an archived edition, kept as published. Its figures were measured on the dates shown and have not been refreshed. Read the current edition

Platform due diligence

Boom Leverage

Platform due diligence dossier

Edition
01
Published
Supersedes
First edition

Companies served

916

prev. 983

Scope: Served domains, FY2021+

Not a decline. The earlier figure counted every ticker in the index, including index-panel seats and issuers whose only filings sit in an era no tier can search. The definition was narrowed to what is actually served.

[1]

Traceable passages

479,663

measured

Scope: Served domains, FY2021+

[2]

Citation gate — Core

98.83%

measured

Scope: FY2021–FY2026

A second, wider measurement exists (99.14%, whole index, 2026-08-14). The two are not averaged and not interchangeable — see appendix rows 3 and 4.

[3]

Markets

1 live · 1 in build

measured

Scope: Thailand serving; Vietnam at corpus stage, coverage zero

[6]

How this document is produced

This dossier is drafted by a language model from a fixed prompt and a supplied evidence pack, then reviewed and published by the founder, who is accountable for every claim in it. It is not an independent third-party opinion and does not pretend to be one. What makes it checkable is the appendix: every figure below carries its scope, the date it was measured, its source, and the way you can re-derive it yourself. Where the evidence pack could not answer a question the structure asks, the question is listed as unanswered rather than filled in.

Questions this edition could not answer

The structure of this dossier asks for the following and the evidence pack did not contain it. They are listed rather than filled in.

  • Customer-side data isolation is described here only at the level this dossier can evidence — market and tenant scoping of the served index. A formal tenancy model, retention schedule and sub-processor list are not published, and this edition does not claim one exists in the form an institutional security review would ask for.
  • No revenue, customer-count, retention or concentration figures appear in this edition. They were not in the evidence pack, and a growth narrative assembled without them would be the one part of this document a reader could not check.
  • Retrieval quality is reported as citation traceability, which measures whether a returned passage can be tied to its page in the filing. It is not a precision/recall benchmark against a labelled answer set, and it should not be read as one.

§1

Executive summary

  • Boom Leverage is qualitative data infrastructure for Asian capital markets: it turns the prose of listed-company regulatory filings — management's own discussion, the risk factors, the auditor's report — into a searchable layer. Live in one market, over 916 companies and 479,663 traceable passages.[1][2]
  • The defensible claim is traceability, not artificial intelligence. 98.83% of passages in the Core Coverage window resolve to the page of the filing they came from; the ones that fail are quarantined rather than displayed, and the quarantine count is published beside the pass rate.[3][10]
  • The commercial model puts two prices on one page in a fixed order: a bespoke report at $1,000 above a subscription that runs the same engine across the whole market daily for a fraction of it.[7][8]
  • Expansion is a franchise, not a rebuild — one engine, one login, one billing account, market-addressed storefronts. The second market is at corpus-acquisition stage with coverage of zero, and its pages say so rather than selling ahead of the index.[6]
  • The concentrations are real and are not softened in §10: one founder, one revenue-generating market, one document source.[11]

§2

What changed since the last edition

This is the first edition, so there is no prior version to diff against. The table below is the dated record of what moved in the weeks before publication.

ItemPreviouslyNowDated
Citation gate, Core Coveragenot published98.83% (353,419 / 357,598)
Companies-served definition983 (every ticker in the index)916 (served domains, FY2021+)
Storefront languageThaiEnglish, with the binding legal texts kept in Thai
Second marketnoneVietnam, at corpus-acquisition stage

§3

The market gap

Every listed Thai company files a One Report (Form 56-1) once a year. The financial statements inside it are structured, standardised, and sold by every data vendor in the region. The prose is not. Management's discussion of what actually happened, the risk factors they chose to name, the matters the auditor decided were key — that material arrives as a several-hundred-page PDF, in Thai, once a year, from roughly nine hundred issuers at the same time.

This produces two different readers with the same problem. A foreign fund cannot read the document at all. A domestic desk can read any one of them and cannot read all of them, which is the same constraint expressed as arithmetic: the analyst covering thirty names has perhaps a day per name per year to spend on the part of the filing that is not a number.

So the qualitative half of the disclosure regime — the half where a deteriorating business explains itself before the numbers turn — is functionally unread at market scale. Not unavailable: unread. That distinction matters commercially, because it means the product is not selling access to documents that are already public and free. It is selling the ability to ask one question across every one of them at once, and to be handed the sentence that answers it, in the company's own words, with the page it sits on.

§4

Extraction architecture and the bounding box

A result in this system is not a passage. It is a passage plus the coordinates that prove where it came from: the filing, the page, and the verbatim sentence as the company wrote it. That bundle is what the term bounding box refers to here — the boundary drawn around a claim so that the claim and its evidence cannot be separated downstream. A retrieval layer that returns text without it is returning an assertion; the same layer with it returns a citation. The index also retains delisted issuers — 860 of the 916 are still listed with the regulator, and the remainder are precisely the companies a failure study needs.[5]

The bundle is enforced by a gate that runs over the index and reports two numbers, not one: how many passages resolved to their page, and how many did not. In the Core Coverage window that is 98.83%, decomposed by axis rather than blended — auditor 99.48%, financials 98.77%, risk 98.85%.[3] A blended figure that cannot decompose is not auditable, and a pass rate published without its quarantine count is a pass rate with the failures deleted.[10]

There is a second measurement over the whole index, taken a day later, of 99.14%.[4] Both are correct. They count different populations on different dates and this document does not average them, reconcile them, or quote whichever is higher — a discipline that is easier to write about than to hold, and the single most common way a metric on a vendor's website becomes untrue without anyone lying.

For a reader whose desk does not read Thai, each cited line arrives twice: the company's Thai sentence verbatim, which is the authoritative text, and an English rendering beside it, labelled as ours rather than as the regulator's. The page number and the link to the document sit on the same row. That construction is deliberate and its limit is deliberate too — the regulator publishes English versions for some issuers and not others, so the product does not promise English filings it does not control.

§5

The false-positive philosophy

The governing assumption is that in this product a wrong answer costs more than a missing one, and the two are not symmetric in any way that a benchmark score captures. A miss costs an analyst a search. A confident false positive costs them a paragraph in an investment committee memo, and it is discovered by the person they were trying to persuade.

That assumption is expressed in four mechanisms rather than in a policy document. Rows that fail verification against their source are withheld and never rendered, so the reader's failure mode is an empty result rather than an unsupported one.[10] Pipeline verdicts combine worst-of and never average, so a category that is 10% complete cannot be carried to green by a sibling that is 90% — the average of a truth and a falsehood is a falsehood with a smaller error bar. Nothing defaults to green: a stage whose status cannot be read fails closed. And scope floors are pinned as literals in code rather than read from configuration, so extending the corpus backwards cannot move a headline coverage figure that was measured on a narrower window.[9]

The same instinct governs the vocabulary. The axis a marketing department would call forensic is labelled operating quality everywhere a customer reads it, because in finance forensic denotes an investigation into fraud, and what the axis actually holds is six lenses over what management wrote about itself. The English term survives in the machine-read surfaces a compliance buyer searches, and nowhere a customer would mistake it for an accusation.

§6

Data integrity, isolation and leak prevention

Three classes of leak are engineered against, and it is worth naming them precisely because only one of them is what the phrase normally means.

The first is cross-market leakage. The engine serves several storefronts from one codebase, so shared navigation, canonical tags and the command-palette search index are all capable of handing a reader of one market the content of another. This is checked by driving a real browser over every page of each market and asserting that no link, no canonical URL and no meta tag crosses the boundary — a source-code grep cannot do it, because the navigation is assembled at render time and a grep would report the defaults as the answer. The failure it exists for is not hypothetical; it shipped once, and was found by a person opening the page.

The second is fabrication leakage: model output that reads like the document but is not in it. The citation gate is the control, and its design decision is that failures are quarantined and counted in public rather than dropped silently.[10] A gate whose rejections are invisible cannot be audited by anyone outside the company, which makes it a claim rather than a control.

The third is staleness leakage — a figure that was true in one month presented as current in another. The build refuses a live figure that jumps beyond a set tolerance and requires a person to record why before the new number may be published, and figures that are deliberately frozen to a measurement date are labelled as frozen rather than refreshed to look current. This is the least glamorous of the three and has caused the most real-world error.

What this edition does not claim is set out in the unanswered questions above: a formal tenancy model, retention schedule and sub-processor list are not published here, and an institutional security review should expect to ask for them separately.

§7

Business model and unit economics

Two prices appear on one page, and their order is the mechanism rather than a layout choice. A bespoke forensic report is priced at $1,000 as a floor, sold per unit, and rendered above the subscription ladder.[7] A reader who meets that number first reads a monthly subscription underneath it as small; a reader who meets the ladder first reads $1,000 as expensive and stops. Anchoring only works in one direction, and the page says out loud why both numbers are on it — an unexplained expensive option beside a cheap one reads as a bait price to exactly the audience being addressed.

Subscriptions are collected in Thai baht. The dollar figures on the site are a reference line derived from a single fixed constant, rounded up, and never from a live exchange rate: a cached marketing page and a checkout page quoting different rates would put the difference on the customer.[8]

Access is metered in credits, and the design test applied to every proposed feature is whether it returns value for the credit it consumes. A feature that burns credits without returning anything produces churn dressed as usage, and a usage metric that rises while satisfaction falls is the most expensive kind of dashboard. Upsell is intended to arrive as a consequence of a customer hitting a ceiling they wanted to hit, not as a prompt.

The bespoke report is also the funnel's top: it is the same memo discipline as this document, performed on a name the client chooses, which makes free research the demonstration rather than the advertisement of the paid kind.

§8

Founder and institutional credibility

The company was founded by Varanchai Yingkhamnueng, previously a Model Risk Manager and Senior Data Scientist inside Thai banking, where the work included an NLP early-warning system for credit risk. He holds an IC Complex Type-1 licence from the Thai SEC.[11]

The relevance is not biographical. Model risk management is the specific discipline of asking what a model does when it is wrong, who finds out, and how fast — and every design decision in §5 is that question applied to a search product rather than to a credit scorecard. Withholding unverifiable rows, combining verdicts worst-of, failing closed on an unreadable status, pinning scope floors so a backfill cannot flatter a headline: these are validation-function instincts, not retrieval-engineering ones, and they are the reason this product is unusually willing to return nothing.

The same fact is a risk, and it appears in §10 in that form rather than being left as an implication here. A methodology that lives substantially in one person's judgement is a methodology with a single point of failure, and an institutional buyer is entitled to price that.

§9

Expansion thesis

The expansion model is a franchise: one engine, one login, one billing account and one set of legal documents, with everything else addressed to a market. Adding a country is a registry row and a tenant configuration file rather than a second product, and the storefront, the coverage dashboard and the nav derive themselves from that row.

Thailand is serving. Vietnam is at corpus-acquisition stage with coverage of zero, and this is the part worth weighing: the Vietnamese storefront publishes that fact, charges nothing, and offers a waitlist instead of a subscription.[6] A greenfield market that were quietly being sold ahead of its index would be the fastest available way to convert the credibility described in §4 into a refund queue.

The condition for opening a market is therefore not commercial appetite but three technical facts: that the raw documents can be obtained at scale and lawfully, that they survive extraction into passages a gate can verify, and that the gate passes at the same bar the incumbent market is held to. A market opened below that bar does not simply underperform — it imports a lower standard into the one asset the brand has, which is the standard itself.

§10

Risk register and open questions

Each entry names a mechanism rather than a category, and states what would show the concern to be answered. A risk with no falsifier is a worry, and a register of worries is decoration.

  1. 01Key-person concentration

    Mechanism
    The extraction methodology, the verification thresholds and the editorial standard all live substantially in one person's judgement. An absence does not degrade the service gradually; it stops the part of it that decides what is allowed to render.
    What would answer it
    A second person with commit rights to the extraction pipeline and a written, exercised handover of the verification standard.
  2. 02Single-market revenue concentration

    Mechanism
    All revenue depends on one market's disclosure regime, one regulator's publication behaviour and one country's institutional buying cycle. The second market has coverage of zero and therefore contributes nothing to diversification today.[6]
    What would answer it
    A first paying customer in a second market, on the same product, at the same verification bar.
  3. 03Dependence on a single document source

    Mechanism
    The corpus is acquired from one regulator's disclosure site under deliberately conservative crawl behaviour. A change to that site's structure, its access controls or its publication format stalls ingestion, and freshness — not volume — is what a subscription is renewed for.
    What would answer it
    A documented second acquisition path, and a published freshness figure that a customer can watch rather than be told about.
  4. 04Translation liability on cited lines

    Mechanism
    The English rendering beside each Thai citation is ours, not the regulator's. A mistranslation on a cited line is the single failure most damaging to a product sold on traceability, because it survives every control that checks whether the passage exists — the passage does exist, and the reading of it is wrong.
    What would answer it
    A sampled bilingual review of cited translations with a published error rate, repeated on a schedule.
  5. 05Metric legibility

    Mechanism
    The platform correctly publishes several coverage and verification figures at different scopes and dates.[3][4][9] A reader who quotes one without its scope misrepresents the company in their own committee memo, and the error will be attributed to us rather than to the quotation.
    What would answer it
    Every published figure carrying its scope inline at the point of quotation — the construction this page is itself an attempt at.
  6. 06Institutional procurement gap

    Mechanism
    The controls described in §6 are engineering controls with automated enforcement. They are not the artefacts an institutional vendor review asks for — a tenancy model, a retention schedule, a sub-processor list, an answered security questionnaire — and a buyer who cannot obtain those may be unable to approve the purchase however good the product is.
    What would answer it
    A completed standard vendor security questionnaire available on request.

§11

Verification appendix

Every figure used above, once, with the population it counts, the date it was measured, its source, and how a third party re-checks it.

#FigurePopulation countedMeasuredSourceHow to re-check it
1916 companies servedDomains actually served, inside every tier's fiscal window (FY2021+), non-companies removedGET /universes (`count`), mirrored as the build-time fallback in src/lib/mdna-stats.tsThe landing page renders this figure from the live engine at build time, so the published claim and the served index cannot disagree without the build failing.
2479,663 traceable passagesRows in the served domains, scoped to FY2021+ — the same window as the 916GET /universes (`rows`)The wider figure (every fiscal year the index holds) is published separately as `rows_searchable`; the two are never added.
398.83% citation-gate pass — 353,419 of 357,598Core Coverage only: FY2021–FY2026, the six years every tier can searchdata/index/th/.citation_gate.json, stamp 2026-08-13T02:40:18ZComponent rates are published with it: auditor 99.48%, financials 98.77%, risk 98.85%. A blended figure that does not decompose is not auditable.
499.14% citation-gate pass — 359,413 of 362,524, with 3,111 quarantinedThe whole served index, all eras — a WIDER population than fact 3, measured a day laterSame gate, run at wider scope⛔ Not reconcilable with fact 3 and not intended to be. Different denominators, different dates. Anyone quoting one of these must quote its scope with it.
5860 of the 916 are still listed with the regulator todayThe remainder are delisted issuers whose filings remain in the indexUniverse build (scripts/universe_build.py)Delisted names are kept on purpose — a failure study needs the companies that failed. Coverage claims count 916; liveness claims count 860.
6One market serving, one in buildThailand (SET/mai) is live. Vietnam (HOSE/HNX/UPCoM) is at corpus-acquisition stage; its coverage is zero and its pages say sosrc/lib/markets.ts (`stage`), and the /vn storefrontThe Vietnamese storefront charges nothing and offers a waitlist rather than a subscription. A market that were secretly live would be selling.
7$1,000 per bespoke report — a floor, not a quoteOne company or one question. Wider scope is quoted individuallysrc/lib/research-service.tsThe per-report price is deliberately not modelled as a subscription tier; running it through the annual toggle would print an annual figure nobody is offered.
8Subscriptions are collected in Thai baht; USD is a reference line onlyOne fixed conversion constant, rounded up, applied site-widesrc/lib/billing.ts (`USD_THB`, `usdRef`); stated in the Terms, clause 6A live FX rate was rejected: a cached page and a checkout page would then disagree, and the difference would land on the customer.
9Archive eras (FY2016–FY2020) are reported separately from Core Coverage, alwaysThe Core window floor is pinned as a literal, not read from configurationpipeline_snapshot_push.py (`SCOPE_MIN`), enforced by tests/test_news_board_era_split.pyBackfilling older years therefore cannot move the headline coverage figure. Measured 2026-08-10: 84.59% including archive, 95.59% excluding, with an identical verified count — the same numerator, two denominators.
10Rows that fail verification against the source are withheld, never displayedAll served axesCitation gate; quarantine count published alongside the pass rate (fact 4)The quarantine figure is published rather than absorbed. A gate whose rejections are invisible cannot be audited by anyone outside the company.
11Founder holds an IC Complex Type-1 licence (Thai SEC)Individual licence; prior role was Model Risk Manager / Senior Data Scientist in Thai banking/th/about/founderLicence status is verifiable through the Thai SEC's public register.

Bespoke Forensic Report

The same discipline, on a name you choose

This dossier is the format applied to ourselves. Commissioned on a company or a question, it is the same memo: findings first, every assertion cited to the filing, page and passage, and the evidence set exported so your analysts can re-run the work instead of trusting it.