Boom Leverage
Platform due diligence

Boom Leverage

Platform due diligence dossier

Edition
02
Published

Companies indexed

916

prev. 916

Scope: Listed Thai companies in the search index, FY2021+

[1]

Citation gate

98.94%

prev. 98.83%

Scope: 371,725 of 375,693 findings, across the Results & financial position, Risk factors and Auditor opinion axes

Same three axes as edition 01, measured three weeks later over a larger corpus — a later reading, not a restatement. Edition 01's 98.83% stands unchanged at its own date.

[2]

Findings withheld

3,968

measured

Scope: Failed the gate and were never rendered — the same population as the pass rate beside it

Published deliberately. A pass rate whose rejections are invisible cannot be audited by anyone outside the company, which makes it a claim rather than a control.

[2]

Markets

1 live · 1 in build

prev. 1 live · 1 in build

Scope: Thailand serving; Vietnam infrastructure live with zero companies indexed

[3]

How this document is produced

This dossier is drafted by a language model from a fixed prompt and a supplied evidence pack, then reviewed and published by the founder, who is accountable for every claim in it. It is not an independent third-party opinion and does not pretend to be one. What makes it checkable is the appendix: every figure below carries its scope, the date it was measured, its source, and the way you can re-derive it yourself. Where the evidence pack could not answer a question the structure asks, the question is listed as unanswered rather than filled in.

Questions this edition could not answer

The structure of this dossier asks for the following and the evidence pack did not contain it. They are listed rather than filled in.

  • No external validation of security or compliance is evidenced. It is unknown whether the platform holds SOC 2 Type II or ISO 27001, and cloud-level tenant isolation for the ALPHA desk licence — a dedicated VPC, for instance — is undocumented. An institutional vendor review should expect to ask for these separately and should not read the architectural controls in §6 as a substitute.
  • No revenue figures appear in this edition: annual recurring revenue, customer count, retention and account concentration were not in the evidence pack. A growth narrative assembled without them would be the one part of this document a reader could not check.
  • The marginal compute cost of a query is not measured here, so no gross margin per credit can be stated. The credit prices in §7 are what a customer pays, not what an action costs to serve — and the two are the whole of the unit economics.

§1

Executive summary

  • Boom Leverage is a qualitative data infrastructure layer for the Thai capital markets. It extracts and indexes the unstructured prose of mandatory regulatory filings — the Form 56-1 One Report, the management discussion and analysis, and the auditor's report — into a searchable semantic vector space covering 916 listed companies.[1]
  • The core technological differentiator is an extraction architecture that pins every result to a physical page coordinate on the source document, governed by a citation gate running at a 98.94% verification pass rate. Anything that fails to map verbatim to the source PDF is withheld, and the 3,968 withheld findings are published rather than absorbed.[2]
  • The commercial thesis converts retail-level keyword searching into institutional checklist scanning. Revenue comes from a metered subscription ladder from a free evaluation tier to a ฿9,900 desk licence, supplemented by bespoke forensic reports at $1,000 each.[5][7][8]
  • The primary vulnerability is extreme key-person concentration. The founder is a single point of failure for the technical architecture, the risk taxonomy and the market expansion strategy, and §10 does not soften it.
  • Expansion into Vietnam is in an unmonetized corpus-acquisition stage: the infrastructure and the price ladder are live, and exactly zero companies are searchable.[3][12]

§2

What changed since the last edition

The period covers a stabilisation of the Thai indexing pipeline and the first infrastructure for regional expansion. The auditor-opinion axis moved from internal testing to public beta; a deterministic rule fix suppressed regulatory boilerplate that had been inflating going-concern alerts; and the Vietnam market went from unannounced to an early-access corpus phase with no filing data available to end users yet.

ItemPreviouslyNowDated
Auditor opinion axisInternal testingBeta — live on the DELTA and GAMMA plans
Boilerplate quarantineRaw wording extractionTSA 570 templates suppressed
Vietnam expansionUnannouncedEarly-access corpus building phase

§3

The market gap

The disclosure regime obliges every listed entity to file extensive qualitative documentation each year. For the Stock Exchange of Thailand, the Form 56-1 One Report runs to hundreds of pages covering operations, risk factors, corporate governance and management's discussion of results. At market scale the qualitative half of those filings remains structurally unread. The quantitative elements — net profit, balance-sheet ratios — are standardised, digitised and available in milliseconds through incumbent terminals. The prose is buried in unstructured, highly variable PDFs. Retail investors rely on delayed third-party summaries; institutional desks fall back on static keyword search. The bottleneck produces a persistent latency: operational distress signals such as shifting customer concentration, localised supply-chain failure or abnormal receivables ageing are formally disclosed by management quarters before they consolidate into a media narrative or a price correction.

"Unread" rather than "unavailable" is the commercially load-bearing word. The alpha does not come from acquiring proprietary data; it comes from reading existing public data comprehensively. Literal string search fails consistently against corporate prose because management teams describe distress in euphemism and highly variable vocabulary. An analyst searching "customer concentration" will miss a filing that says "the majority of the Company's revenue recognized… comes from government and state enterprises customers", because the exact character string is absent. A system without semantic comprehension does not find the risk.

Boom Leverage occupies that void by converting the whole market's Form 56-1 filings into a semantic vector space, shifting the institutional workflow from guessing keywords to applying a standardised risk taxonomy across the entire cross-section at once.[1] The value is anchored in corporate memory rather than prediction: the platform holds the historical baseline of disclosures up to six years deep, which is what surfaces the exact moment a management team changed its story about a specific operational risk.[6] The gap the product fills is the human inability to track semantic change across thousands of unstructured documents over multi-year periods.

A further 206,758 findings have been verified in the FY2016–FY2020 archive era, which is built and checked but not yet entitled to any tier. It is reported separately from the served window and never added to it — the two have different denominators, and merging them is precisely the flattery this document is written against.[4]

§4

Extraction architecture and the bounding box

The extraction architecture is built on a spatial bounding-box paradigm, enforcing the constraint that a valid result consists of four inseparable components: the literal passage, the specific filing, the exact page number, and the verifiable coordinates of the verbatim sentence on that page. Rather than operating purely in string space, the engine ingests the Form 56-1 PDFs and records spatial metadata. On a query, a multilingual embedding model retrieves the relevant passage; where a document is a scanned image with no embedded text layer, optical character recognition and vision models are deployed strictly to recover word coordinates. The OCR layer is prohibited from altering semantic content. It is an alignment tool and nothing else.

The defining control is the citation gate. Before any result renders, the system runs a deterministic match between the extracted passage and the source text sitting at the recorded PDF coordinates. If that match fails for any reason — formatting corruption, OCR failure, model hallucination — the output is suppressed. "Accurate" is not a functional claim in this architecture; "withheld when the passage cannot be resolved to a page in the filing" is the operable mechanism. As of 4 September 2026 the gate recorded a 98.94% pass rate across the Results & financial position, Risk factors and Auditor opinion axes, and the withheld count — 3,968 findings — is published on the dashboard beside it, which is what makes the failure rate auditable rather than asserted.[2]

For institutional users who do not read the source language, the architecture carries a non-destructive verification path. Every cited line is presented bilingually: the company's original Thai sentence verbatim, holding its status as the authoritative text, and the English rendering beside it, explicitly labelled as a translation and as ours. Because the filing link and the page coordinate arrive on the same row, an English-speaking portfolio manager can escalate any flagged sentence to local counsel or a bilingual analyst and have it settled in minutes. The architecture refuses to compel trust in an unverified translated string.

The verification path is metered rather than free, and the prices state what the system considers expensive: a market-wide search costs one credit, opening a filing page with the quote boxed costs two, and revealing one company's auditor findings costs five — while following the citation out to the regulator's own site costs nothing at all, because at that point the reader is checking us against the source and no compute of ours is involved.[10]

§5

The false-positive philosophy

The platform operates on a philosophy taken directly from bank-grade model risk management: a false positive is more destructive to institutional trust than a missing signal. A system that routinely paints benign disclosure as severe operational risk trains portfolio managers to ignore its output entirely. Boom Leverage therefore deprioritises alert volume in favour of signal density. Where an extracted passage cannot meet the bounding box's verification standard, the system fails closed — it withholds the finding rather than interpolating, estimating or averaging around it — and the withheld count is pinned to the interface, forcing the system's operational limits into view instead of flattering its parsing accuracy.[2]

The discipline is enforced hardest on the auditor-opinion axis. An earlier iteration pushed every identified risk category into the primary alert bucket and produced a median of 240 red flags per company. Because the rules oblige every listed entity to disclose standard liquidity and credit risk, blanket extraction produces unactionable noise. The revised mechanism evaluates the specific lexical wording of the auditor's text rather than triggering on the presence of a mandatory structural header, isolating modified opinions — qualified, adverse, disclaimer — from routine Emphasis of Matter or Key Audit Matters. The direct consequence of that threshold: out of a scanned 910-company archive, 68 entities were flagged.[9][11]

The system also quarantines regulatory boilerplate so that a structural artefact cannot present as an idiosyncratic risk. In an August 2026 pipeline update the same logic was applied to the TSA 570 going-concern template, where the engine had been flagging the standardised sentence "However, future events or conditions may cause the Group and the Company to cease to continue as a going concern" as an acute material uncertainty. A deterministic rule fix quarantined 96 boilerplate rows and removed 60 otherwise clean companies from the red-flag register, while the 68 modified opinions did not move — which is the check that matters, because a filter that also deletes true positives is not a fix.[11]

The philosophy's purpose is that when the system does raise an alert, it represents a deliberate disclosure of distress by management or the auditor rather than a technical artefact of the parsing taxonomy.

§6

Data integrity, isolation and leak prevention

Data integrity rests on the immutability of the source regulatory text; isolation prevents cross-tenant contamination. The primary leak class in semantic vector infrastructure is query visibility — where one institutional tenant's search terminology informs the autocomplete, caching layer or model weights that another tenant meets. The platform enforces isolation by keeping a rigid separation between the static, market-public index and the dynamic query layer: user search histories and applied taxonomies are not aggregated into a centralised continuous-learning dataset. The mechanism exists so that a proprietary risk hypothesis developed by one desk is not broadcast to the market through a fine-tuned model update.

The second leak class is generative hallucination, typically reached through prompt injection. It is mitigated by removing generative synthesis from the core search loop entirely: the architecture is restricted to vector retrieval and coordinate mapping, and the interface holds no generative mechanism authorised to summarise, infer or forecast. Adversarial input is processed as a literal semantic vector, matched against conceptually proximate paragraphs, and put through the bounding-box citation gate. Where no verbatim match exists on a PDF page, the query fails safely and returns nothing.

The third class is staleness — a figure that was true in one month presented as current in another. Figures frozen to a measurement date are labelled as frozen rather than refreshed to look current, which is the rule this dossier applies to itself: edition 01's 98.83% was not overwritten by this edition's 98.94%, and the two remain readable side by side with their own dates and populations.[2]

What is not evidenced is set out in the unanswered questions above. There is no external attestation here — no SOC 2 Type II, no ISO 27001, no documented cloud-level tenant isolation for the desk licence — and the architectural controls described in this section are not a substitute for one. An institutional vendor review should treat that gap as open.

§7

Business model and unit economics

The model is a metered freemium subscription ladder designed to remove evaluation friction before converting to recurring revenue. The entry point is a Google sign-in granting 10 free credits a day with no payment details.[13] DELTA sits at ฿799 a month for four years of history and 150 credits a day, and is the intended conversion point; GAMMA at ฿1,399 buys 350 credits and the full six-year served window.[5][6] Prices are denominated in baht rather than dollars, which removes FX volatility from a regional broker's or asset manager's operating budget and keeps the page and the checkout quoting the same number.

Usage is metered by action rather than by seat, and the price list is the product's opinion about where its value sits: one credit for a market-wide search, two to open a filing page with the quote boxed, five to reveal one company's auditor findings.[10] The dearest action is the verification, not the search. Upsell is mechanical rather than prompted: an analyst on the free tier who finds a distress signal in a current filing exhausts the daily quota tracing that narrative backwards through the archive, and the constraint that bites is history depth — which is what DELTA and GAMMA sell.

At the top of the ladder the ALPHA desk licence starts at ฿9,900 a month for multiple seats, API access and bulk CSV/Excel evidence export.[7] The ladder is anchored against a bespoke consulting service — On-Demand Research at $1,000 per report, delivering a fully cited forensic memo, the raw evidence set, and a call with the analyst who wrote it.[8] Presenting the report price alongside the subscription ladder, and above it, is the mechanism: a reader who meets $1,000 first reads ฿799 a month as inexpensive relative to the cost of doing the same diligence by hand.

What this section cannot state is the other half of unit economics. The marginal compute cost of serving a credit is not measured in this edition, so no gross margin per action is claimed here.

§8

Founder and institutional credibility

The platform is founded and operated by Varanchai Yingkhamnueng. His background sits in quantitative model risk management and banking analytics rather than in software engineering: trained in the hard sciences — B.Sc. Chemistry, M.Sc. Physical Chemistry with highest distinction — then quantitative proprietary trading, then institutional banking. At Kasikornbank he worked in credit-risk quantitative analytics and built the institution's first natural-language early-warning system, deep learning over news and prose to catch credit-risk signals before they became non-performing loans. At Krung Thai Bank he was a Senior Data Scientist in business risk and macro research, owning and validating the production monthly expected-credit-loss models the bank used under TFRS 9.

That pedigree dictates the platform's risk posture. A practitioner who has had to defend a TFRS 9 ECL model in front of internal bank auditors acquires a durable distrust of unexplainable variables, and the structural refusal of generative synthesis, the false-positive discipline in §5, and the wording-severity taxonomy on the auditor axis are all that framework translated into a commercial product. The Thai SEC IC Complex Type-1 investment consultant licence anchors the same adherence on the regulatory side.[9]

The same fact is the platform's largest unmitigated risk, and it is stated here rather than left as an implication. The entity runs on an extreme single-key-person dependency: the technical architecture, the continuous tuning of the risk taxonomy, the bilingual prompt engineering and the go-to-market execution rely on one person. The evidence pack documents no engineering team, no co-founder and no institutional backing. Were the founder incapacitated or to pivot away, the platform's ability to maintain the vector index, adapt its parsers to inevitable regulatory format changes, or execute the multi-market thesis would stop.

§9

Expansion thesis

The expansion thesis is a franchise mechanism: replicate the semantic indexing and bounding-box extraction engine across adjacent markets that have a similar centralised, mandatory disclosure framework. Vietnam is the designated second market, targeting the Annual Reports (Báo cáo thường niên) of entities listed on HOSE, HNX and UPCoM. The rationale is lateral scaling — reusing the existing multilingual embeddings and the coordinate-mapping software to parse a new language without redesigning the core search and verification loop.

The Vietnamese integration is strictly at corpus-acquisition stage. The search infrastructure is live alongside the Thai one and the localised ladder is published at ₫639,000 a month for DELTA, derived one-to-one from the Thai price at a fixed rate.[12] Exactly zero Vietnamese companies are currently indexed and searchable.[3] The gating condition is deliberate: no market opens to paid queries until its corpus clears the same verification pass-rate threshold enforced in Thailand.[2] The early-access waitlist lets clients request priority indexing for named companies, which implies demand-driven processing rather than a resource-intensive market-wide ingestion before launch.

Three operating conditions have to hold for the thesis to be commercial. The regulatory formatting of Vietnamese filings must carry a predictable text layer, or remain tractable to the vision models, so the citation bounding box functions. The embedding model must translate opaque Vietnamese business terminology into operational distress signals as accurately as it currently does in Thai. And localised institutional demand must validate the pricing without requiring bespoke product adjustments that would fork the engine. Until the pipeline delivers a verified Vietnamese corpus, the expansion thesis is a capability rather than a business.

§10

Risk register and open questions

Each entry isolates a structural dependency, a regulatory exposure or a technical bottleneck, and states what would show the concern answered. The open questions above — unverified recurring revenue, unmeasured compute cost per query, and the absence of formal SOC 2 or ISO 27001 certification — inform this register rather than sitting apart from it.

  1. 01Key-person concentration

    Mechanism
    The entire technical stack, the specialised banking risk taxonomy and the go-to-market execution rely on a single founder. Incapacitation or a strategic pivot halts pipeline updates, regulatory adaptation and regional expansion immediately.
    What would answer it
    The firm hires a dedicated engineering lead, or secures institutional capital that enforces operational redundancy and documented knowledge transfer.
  2. 02Regulatory format dependency

    Mechanism
    Bounding-box extraction depends on the structural predictability of the Form 56-1. If the regulator materially alters the required taxonomy, or permits disparate XBRL tagging without a rigid prose mandate, the parser stops classifying items into the established lenses.
    What would answer it
    The regulator moves to raw-text APIs instead of rendered PDFs, which removes the spatial-parsing dependency altogether.
  3. 03Translation liability on cited lines

    Mechanism
    The Thai text is preserved as the sole authoritative artefact and the interface says so, but an institutional client who acts on a mistranslated English rendering still takes the loss — and the dispute lands on us. It is the failure that survives every control checking whether the passage exists, because the passage does exist and the reading of it is wrong.
    What would answer it
    A sampled bilingual review of cited translations with a published error rate, repeated on a schedule — the control that would make the exposure measurable instead of merely disclaimed.
  4. 04Vision-model exhaustion on scanned filings

    Mechanism
    Listed entities frequently file image-only scanned PDFs. Recovering precise word coordinates from those with vision models carries both a compute cost and a failure rate, and every failure is a finding withheld — so the same mechanism that protects the reader from a false positive also quietly removes real signal from the market.
    What would answer it
    The citation pass rate across image-heavy filings holds above 98% without compute cost rising disproportionately — reported by filing type rather than as one blended figure.[2]
  5. 05Commoditisation by an incumbent

    Mechanism
    Established terminal providers hold enormous distribution moats. If one deploys equivalent semantic vector search over Thai or Vietnamese filings natively, the technological edge commoditises immediately and what remains is the discipline, not the index.
    What would answer it
    An incumbent ships the feature and its alerts prove too noisy for institutional risk desks, because the false-positive discipline in §5 is a product decision rather than a technical one — which is the only version of this outcome that is a moat rather than a reprieve.

§11

Verification appendix

Every quantitative claim in the prose above, bound to the population it counts, the date it was measured, its source and the way a third party re-checks it. A figure absent from this table is unsupported by the evidence pack and has been kept out of the analysis.

#FigurePopulation countedMeasuredSourceHow to re-check it
1916 companiesListed Thai companies in the search index, FY2021+GET /universes, rendered on the TerminalRun a whole-market query on the Terminal: the universe count it searches is the same figure, read from the served index rather than from this page.
298.94% citation-gate pass — 371,725 of 375,693 verified, 3,968 withheldThe Results & financial position, Risk factors and Auditor opinion axesTerminal instrumentation dashboardNumerator, denominator and withheld count are published together, so the arithmetic closes in public. A pass rate published without its withheld count is a pass rate with the failures deleted.
30 Vietnamese companiesCurrently indexed and searchable on the Vietnam marketVietnam market pageThe page offers a waitlist rather than a subscription and charges nothing. A market being sold ahead of its index would be selling.
4206,758 findings verifiedThe FY2016–FY2020 archive era — built, verified, and not yet entitled to any tierArchive pipeline dashboard⛔ Reported as Archive Coverage and never added to row 2. The Core window floor is pinned in code, so a backfill cannot move the headline figure — different denominators, deliberately kept apart.
5฿799 per month — DELTAFour years of history, 150 credits a dayPricing pageRendered from `MDNA_TIERS` in src/lib/mdna-tiers.ts — the same constant the Stripe price id is looked up from, so the page and checkout cannot disagree.
6฿1,399 per month — GAMMASix years of history — the full served window — 350 credits a dayPricing pageSame constant as row 5. The six years match the Core Coverage window the citation gate in row 2 is measured over.
7฿9,900 per month — ALPHAStarting price for the desk licence: multiple seats, API, bulk CSV/Excel evidence exportPricing pageSame constant as rows 5 and 6. "Starting" is load-bearing: seats beyond the base are quoted.
8$1,000 per reportStarting price for one On-Demand bespoke research report — one company or one question; wider scope is quotedEnterprise page (src/lib/research-service.ts)Deliberately not modelled as a subscription tier: run through the annual toggle it would print a yearly figure nobody is offered.
968 companies of 910 scannedModified auditor opinions — qualified, adverse or disclaimer — across FY2021–FY2025The auditor-opinion research articleThe article quotes each opinion with its page number and a link to the filing at the regulator, so the count can be re-derived name by name rather than believed.
101 / 2 / 5 credits per actionOne market-wide search · opening a filing page with the quote boxed · revealing one company's auditor findings. Following the citation to the regulator's own site costs nothing`MDNA_CREDIT_COST` in the TerminalAll actions draw from one daily bucket, so the meter in the account menu is the whole truth rather than a search counter. The auditor unlock is charged once per company per window.
1196 boilerplate rows quarantined, clearing 60 companies; median before the fix, 240 items per companyThe TSA 570 going-concern template, inside the red-flag bucketThe auditor-opinion research articleThe article publishes the before-and-after table (772 rows → 676) and confirms the 68 modified opinions in row 9 did not move. A filter that also removed true positives would show up there.
12₫639,000 per month — DELTA, VietnamThe Vietnamese ladder's entry package, published while coverage is zeroVietnam pricing (src/lib/vn-pricing.ts)Derived one-to-one from ฿799 at a fixed 1 THB = 800 VND, not from a live rate. Nothing is charged until the corpus is live.
1310 free credits a dayGranted on Google sign-in, no payment details required`MDNA_FREE` in src/lib/mdna-tiers.ts10 free credits against a 5-credit auditor unlock and a 1-credit search is exactly one company unlocked per day, and not two — the free tier is sized to demonstrate the dearest action, not to ration it into uselessness.

Earlier editions

Archived editions are kept verbatim and are not updated. Their figures were correct on the date each was measured and should be quoted with that date attached.

Bespoke Forensic Report

The same discipline, on a name you choose

This dossier is the format applied to ourselves. Commissioned on a company or a question, it is the same memo: findings first, every assertion cited to the filing, page and passage, and the evidence set exported so your analysts can re-run the work instead of trusting it.