You lost the moment you didn't know what to type — institutions don't guess keywords, they run a checklist
Everyone does the same thing when handed a new search box or a new AI: stare at the blinking cursor, then type something broad like 'is this company any good' — get garbage back, and conclude the tool is weak. The problem is the question. Bank risk departments never guess keywords. They sweep the market with a taxonomy: a standing checklist of topics that gets reviewed every quarter. This is how the Scan page puts that taxonomy — 30 topics across 6 groups, plus a 5-topic auditor's report group — into your hands, and why every row that cannot point back to a source page number is killed before it reaches your screen.
You lost the moment you didn't know what to type — institutions don't guess keywords, they run a checklist
Everyone does the same thing when handed a new search box or a new AI: stare at the blinking cursor, then type something broad like 'is this company any good' — get garbage back, and conclude the tool is weak. The problem is the question. Bank risk departments never guess keywords. They sweep the market with a taxonomy: a standing checklist of topics that gets reviewed every quarter. This is how the Scan page puts that taxonomy — 30 topics across 6 groups, plus a 5-topic auditor's report group — into your hands, and why every row that cannot point back to a source page number is killed before it reaches your screen.
You know the moment. New tool, screen open, an empty search box and a cursor blinking at you.
And your head empties out with it.
So you type something broad — "บริษัทนี้ดีไหม" [is this company any good — a question that asks for everything and therefore anchors to nothing] or "แนวโน้มธุรกิจเป็นยังไง" [how is the business trending — same problem, no document, no page, no claim anyone can check] — you hit search, you get back a lump of average-sounding text that tells you nothing, you decide "this tool isn't good enough yet," and you close the tab.
I watched this happen over and over while building an Early Warning System inside a bank, and here is the part nobody says out loud: the people who extract real value from a document search tool are the people who already know what they are hunting for. The people who would gain the most — the ones who don't yet know which risk should worry them — are the ones who can't type anything and quit first.
The tool isn't the problem. Neither is your intelligence. The system hands the hardest job back to you: know roughly what the answer is before you are allowed to ask the question.
1. The institutional secret nobody bothers to mention, because they assume everyone knows it
On the risk side of a bank, no team ever sat around in the morning improvising keywords. Nobody asked "what should we search for today?"
Because the checklist already existed.
The formal name is taxonomy — a standing structure of risk topics swept every cycle whether or not you happen to think of them. What to look at in macro. What to look at in the supply chain. What to look at in capital structure. What to look at in regulation. A list crystallized out of accumulated post-mortems: which channel did the ones that blew up start signalling from?
The power isn't in the cleverness of any single topic. It's that it doesn't let you forget. You may not think about concession renewal once all year. If it sits on the checklist, it gets swept every quarter regardless of your mood that day.
That is the real difference between retail searching and institutional scanning, and none of it is about technology.
| Guessing keywords yourself | Sweeping with a taxonomy (institutional) | |
|---|---|---|
| Starting point | Blank screen, blinking cursor | A checklist that already exists. Click and go. |
| Coverage | Only what you happen to think of that day | Every topic on the list, thought of or not |
| Blind spot | What you don't know to worry about — which is exactly what takes people out | Swept automatically, because it sits on the list |
| Different result every time? | Yes. Depends on which words you typed that day | No. Repeatable, comparable across quarters |
| Quality depends on | The asker's experience and memory | The structure of the list, not the mood of the day |
| Time per pass | Long, because thinking up the words eats it | One click |
Row three is the one that costs money. What wipes people out is almost never the thing they were already watching. It is the thing that never crossed their mind at all, and keyword guessing, by its own definition, can never walk you into your own blind spot.
2. Scan puts that checklist in your hands
Scan is not a convenience shortcut, and it is not a set of canned queries for people who can't be bothered to type. It is the entire taxonomy turned into buttons. You don't compose a query. You pick one topic. The rest is the system's job.
This is the real checklist on the screen, not an illustration:
🌍 Macro & geopolitics War & geopolitics · FX volatility · Policy rate hikes
· Cost-push inflation · Recession, demand shrinks
🚢 Supply chain & operations Chip shortage · Freight & logistics · Customer concentration
· Labor shortage · Raw material & energy costs
📉 Market & competition Price wars · Consumer behavior shifts · Foreign competitors
· Market share loss · New product failure
💰 Financial (Debt load · Liquidity · Receivable quality, etc.)
⚖️ Regulation & ESG Government policy/law · Carbon tax · Community & environment
· Concession/license renewal · Minimum wage
💻 Technology & cyber Cyber & data breach · Disruption · IT systems down
· IT talent shortage · System upgrade costs
─────────────────── Total 30 topics · 6 groups
🔍 Auditor's report [BETA] Going Concern · Qualified opinion · Disclaimer of opinion
· Emphasis of matter · KAM: revenue recognition
↳ Different document from MD&A — see limits at the end
Read that list slowly once, then answer honestly: without it, how many of these would you have come up with? Most people get 4–5, the ones currently in the news. The rest is blind spot, and blind spot doesn't mean unimportant. It means you won't see it until it has already become a lower share price.
The last group deserves its own note, because it comes from a different document. Not management's words — the auditor's report, written by an outside party whose job is to object. Buttons like Going Concern (material uncertainty related to going concern) and qualified opinion are the first thing any credit officer reads, and the thing retail almost never opens. This group is still BETA because the auditor-report library doesn't cover every company yet (see the limits at the end).
3. What happens when you click one topic
Click "Policy rate hikes" once and the system reads the management discussion and analysis (MD&A) in the Form 56-1 (One Report) filings of 916 listed companies, then returns which of them wrote about it themselves — with the actual sentence the company wrote and the page number, one click from the source document.
(If you're not yet sure where the MD&A sits inside Form 56-1 and which section to read first, I mapped the whole filing in How to actually read One Report (56-1 / MD&A): 4 sections, 9 topics.)
The difference is "wrote about it," not "contains the string." The search runs on meaning. A company that wrote "ต้นทุนทางการเงินเพิ่มขึ้นจากภาระดอกเบี้ยเงินกู้" [finance costs went up because the loan book repriced — they never name the central bank] with the phrase "ดอกเบี้ยนโยบาย" [policy rate, the exact words you would have typed] nowhere in the document still gets pulled. That case is the norm, not the exception: nobody writes an annual report in the words you happen to think of.
Why this matters: The slowest part of reading a Form 56-1 is not the reading. It is finding which page of which company to read. Scan deletes that step entirely and leaves you the only work a human has to do — deciding what the company's own sentence actually means.
4. Every sentence must trace back to a source page. Can't trace = doesn't display
This is the part I am most protective of, and the reason I built the tool myself instead of bolting an LLM onto a pile of PDFs and calling it done.
Every row Scan returns clears a citation gate before it reaches your screen, and the gate is blunt:
System finds a sentence matching the topic you clicked
│
▼
Take that sentence and locate it in the source PDF the SEC published
│
┌───────────┴───────────┐
▼ ▼
Position found Not found / page is a scanned image
+ page number no text layer to verify against
│ │
▼ ▼
✅ SHOW 🗑 DROP — cut, never rendered
with page number no matter how high the match score
+ source link
These are not trivial numbers. On one real topic search on 5 August 2026, the system killed 15 rows out of a single result set — not_located on 11 rows, no_text_layer on 3, and no_source_pdf on 1. What you saw was what survived.
Most tools take the opposite route: show everything, look smarter, return more, and nobody catches it. I spent three years in model risk and I know what one wrong number costs once the report has already gone out.
A result you can't trace back has negative value, not zero — because it will talk you into believing it.
Zero means you got nothing, which is fine; you go look somewhere else. Negative means you got something credible enough to act on with no way of knowing whether it is true, and you find out when it is already sitting in a published report, or in an order you already filled.
Why this matters: If you put this into a research note, in front of a client, or on a board table, you have to point back at the source page in one click. Scan is built so the question "where did this come from" always has an answer — because if it doesn't, it had no business being on the screen in the first place. (The same standard regulators enforce on AI inside financial institutions — the Bank of Thailand's ban on black-box AI.)
5. The AI reads for you. It doesn't write on the company's behalf
Scan does not summarize for the company and does not interpret. What you get is the sentence the company wrote, lifted verbatim.
That is a deliberate decision, not a technical limitation. The moment a system starts "summarizing for you," it becomes the author, and you start believing what the model invented without noticing. I've already written up the mechanism that lets fabricated numbers escape in Why AI invents numbers, and the 3 gates I use to stop it — Scan is those gates shipped as a product.
The AI's job here ends at "find it." Interpretation is still yours, and it should be.
6. Limits worth knowing before you use it
I'm writing this section because a tool that won't state its own boundaries is the tool that burns you on the day you trust it most.
- Coverage runs FY2021–2026. Comparisons further back than that are not in the system yet.
- Quoted text comes back in English, because the index is built from the English-language versions companies file. The system doesn't translate, for the reason in section 5.
- The auditor's report group is still BETA. Every passage clears the same citation gate as every other group, but the library doesn't cover every company yet. Filings whose auditor's report hasn't reached us won't appear in results — so "not found" in this group does not yet let you conclude the auditor raised nothing.
- Some companies file documents text can't be extracted from (scanned to image, no text layer). Those rows are cut by the citation gate under section 4, which means "not found" doesn't always mean "the company didn't write it."
- Scan answers who wrote about what. It does not answer whether to buy or sell, and it won't.
7. Try it yourself. No card
Scan is live at terminal.boomleverage.com/scan — a free account gets 10 runs a day, no card required, enough to walk every position you hold inside a week. For the launch, anyone who logs in gets the DELTA pack (4 years of history) free until 30 September 2026, also without a card (full terms).
If you're a team that needs bulk pulls, file exports, or a connection into your own systems, look at the enterprise plans.
If you want to see what one topic swept across the whole market actually returns before you try it yourself, I published the full set in Who wages and rates are biting: 23 companies, verbatim from Form 56-1 — the output of firing exactly two queries into this same library.
How I'd use it the first time, and it runs against instinct: don't click the topic you're already worried about. Click the one you have never once considered for that stock. The topic you're worried about, you watch every day, and so does the market. The topic that has never crossed your mind is the only channel with anything left in it.
That is why the checklist exists at all — not to save you typing, but so you never have to rely on your own memory on the day it matters most.
Read next
One Red Flag Was Never Enough: I Counted 104,153 Lines Thai Companies Wrote About Themselves, and All Six Lenses Lit Together Only 9 Times
Read more MD&AFunds Never Read a 56-1 From Page One — Here Is the Playbook for Where They Actually Open (a 4-Part, 9-Item Map + the 6 Places Everyone Skips)
Read more MD&AEarnings Can Be Dressed Up. The Cash Has to Land Somewhere — 216 Disclosures Companies Wrote Themselves, on Pages Nobody Reads
Read more