Boom Leverage
All insights

Ctrl+F Doesn't Just Miss — It Finds the Wrong Thing, and on Some Files It Sees Nothing at All: Three Layers of Information Investors Lose in Thai Filings

Everyone knows Ctrl+F misses the euphemism. Fewer people know it also floods you with the right word in the wrong meaning — and almost nobody checks whether the PDF has any text in it at all. I opened the filings to count: 22.7% of Thai MD&A filings from FY2016–FY2020 (2,945 of 12,948) carry no searchable text layer, and in a 3,000-file sample from FY2021–FY2026 it is still 14.0%. This piece walks a real 18-page CENTEL filing where Ctrl+F finds zero characters, shows the OCR errors that make a naive fix dangerous, gives you a 10-second test for any PDF, and explains the rule the Terminal runs on: every result points back to the exact page, and anything that cannot be traced is dropped.

Varanchai Yingkhamnueng·
MD&ABoom Leverage

Ctrl+F Doesn't Just Miss — It Finds the Wrong Thing, and on Some Files It Sees Nothing at All: Three Layers of Information Investors Lose in Thai Filings

Everyone knows Ctrl+F misses the euphemism. Fewer people know it also floods you with the right word in the wrong meaning — and almost nobody checks whether the PDF has any text in it at all. I opened the filings to count: 22.7% of Thai MD&A filings from FY2016–FY2020 (2,945 of 12,948) carry no searchable text layer, and in a 3,000-file sample from FY2021–FY2026 it is still 14.0%. This piece walks a real 18-page CENTEL filing where Ctrl+F finds zero characters, shows the OCR errors that make a naive fix dangerous, gives you a 10-second test for any PDF, and explains the rule the Terminal runs on: every result points back to the exact page, and anything that cannot be traced is dropped.

CENTEL · MD&A Q2/2020 · page 15: "Liquidity management : The Company has requested commercial banks for additional credit facilities relating to both short term and long term loans totaling Baht 5.5 billion; whereby Baht 4.0 billion has yet still to be drawn down, together with cash / cash equivalents on hand totaling Baht 2.9 billion being also available for use as at June 30, 2020. This should be sufficient to support any operating expenses in the event that the ongoing COVID-19 pandemic situation has not improved until the mid of next year."

— verbatim from the filing, page 15 of 18 · company filings at the SEC · the PDF itself

That paragraph is the single most useful thing a hotel operator could tell you in the middle of 2020: how much funding it had lined up, how much was still undrawn, and how long management thought the cushion would last.

Now download that PDF and press Ctrl+F. Type liquidity. No matches. Type COVID — a word that appears on these pages at least ten times. No matches. Type Baht. No matches.

Your PDF reader is not broken, and you did not misspell anything. The file contains zero characters. All 18 pages are pictures of paper. To your search box, this filing is blank — and it looks exactly like every other PDF on your screen, so nothing warns you.

I want to be precise about what this article is not. It is not a story about CENTEL: management disclosed a funding buffer in plain language, which is exactly what a filer should do. It is a story about where that disclosure was unreachable, and how often the same thing happens across the market.

What I measured, and its limits: Every number below comes from opening the PDFs directly on 13 September 2026. "No searchable text" uses the same rule our OCR pipeline uses: fewer than 200 extractable characters across the first 8 pages. That rule is conservative — a file that is half digital and half scanned passes as "searchable", so the real blind share is, if anything, higher. FY2016–FY2020 MD&A is a full count (12,948 files). FY2021–FY2026 MD&A is a seeded random sample of 3,000 files out of 18,815, so treat those percentages as ±1–2 points. One Reports are a full count of the 2,868 in our archive.

1. Ctrl+F fails in three layers, not one

When people complain about keyword search, they almost always mean the first layer. The other two do more damage, because you do not notice them happening.

LayerWhat goes wrongWhat it feels likeWhy it hurts
1. Different words, same meaningManagement writes around the word you typed"No matches"You conclude the risk is not there
2. Same word, different meaningThe word matches, the meaning does not"199 matches"You drown, skim, and stop reading
3. No words at allThe PDF is an image with no text layer"No matches" — identical to layer 1You conclude the risk is not there, on a file nobody could have searched

Layer 1 is the euphemism problem — nobody writes "we are running short of cash", they write "prudent cash-flow management". I covered it in depth, with three live cases, in why Ctrl+F has always lost retail the trade, so I will not repeat it here.

This article is about layers 2 and 3, and mostly about layer 3 — because layer 3 is the one you cannot fix by being a smarter searcher.

Why this matters: layers 1 and 3 produce the same screen: "No matches found." One means "try a different word". The other means "there is nothing here for any word to match". A reader who cannot tell them apart will treat a blank search on an image as evidence that the risk does not exist.

2. Layer 2: the word matches and the meaning doesn't

Say you want to know how much a company borrows. The obvious search is loan. Here is what that word means in three real filings, each quoted exactly as written:

TKS · MD&A Q1/2024 · page 5: "Net cash flows provided by investing activities of THB 57.5 million, mainly from the Company paid for the purchase of fixed assets of THB 11.8 million. While there was a cash received from sold warrant of THB 5.5 million and repayment from loan to employees of THB 4.6 million."

— verbatim · original at the SEC

Here "loan" is money owed to the company by its own staff, coming back in. It is an asset being repaid, not debt.

FSMART · MD&A Q3/2025 · page 2: "The Company focuses on offering loans to employees of large organizations, with a membership base of over one million individuals."

— verbatim · original at the SEC

Here "loan" is the product. It is how the business earns interest income, not a liability on its balance sheet.

BYD · One Report FY2024 · page 273: "The Company has established a labor protection and welfare committee to advise and suggest opinions to employers on providing appropriate welfare to employees, including provident funds, loans to employees, and life and health insurance, organizing activities, recreation, and training to provide knowledge."

— verbatim · original at the SEC

Here "loan" is a staff benefit sitting in the sustainability section.

Three companies, one word, and not one of these is the interest-bearing debt you were looking for. Now scale it up. In that same BYD One Report, the string "loan" appears 199 times across 62 of its 304 pages. Somewhere in those 62 pages is the borrowing you actually care about. Ctrl+F gives you all 199 in the order they appear, with no idea which is which.

The same trap sits in Thai-language filings. Search เงินกู้ [loan] and you will also land on phrases like เงินให้กู้ยืมแก่บริษัทย่อย [loans to subsidiaries], which turns up in management commentary across the market — money the company lent out, the opposite direction from what you were trying to measure.

A rule from risk work: a match count measures how often a word appears, not how much a thing matters. In credit review, nobody reads a borrower's file by counting keywords. You read the sentence the word lives in. A tool that cannot tell "we owe" from "we are owed" is not a search tool for credit questions.

3. Layer 3: the file where Ctrl+F is blind

Back to CENTEL. The reason the hook paragraph was unreachable is not the wording. It is the file format.

A PDF can carry two things: the picture of the page, and a layer of text that sits behind the picture. Ctrl+F, copy-paste, screen readers and every search engine only ever touch the text layer. When a company prints a document, signs it, and scans it back in, the result is a PDF with the picture and no text layer. On screen it looks identical. To every machine, it is empty.

So I counted how often that happens in the filings that matter most for reading management's own narrative: the MD&A (management discussion and analysis), which Thai listed companies file every quarter and at year-end.

Fiscal yearMD&A files checkedNo searchable textShare
FY20162,37455223.3%
FY20172,44658624.0%
FY20182,54261724.3%
FY20192,69956821.0%
FY20202,88762221.5%
FY2016–FY2020 (full count)12,9482,94522.7%
FY2021 (sample)4759018.9%
FY2022 (sample)5337714.4%
FY2023 (sample)5738915.5%
FY2024 (sample)5476411.7%
FY2025 (sample)6076610.9%
FY2026 (sample, year in progress)2653312.5%
FY2021–FY2026 (3,000-file sample)3,00041914.0%

Read that table two ways.

Going back in time, it gets much worse. Between FY2016 and FY2020, roughly one MD&A in four is an image. If you are building a ten-year track record of what a management team said — which is exactly the kind of reading that separates a durable business from a lucky one — about a quarter of your source pile cannot be searched at all.

It has not gone away. It has improved, but in FY2024–FY2026 it still sits around one file in eight to one in nine. This is not a relic of the 2010s.

Now compare the other two document types I checked:

DocumentFiles checkedNo searchable text
One Report (Form 56-1), FY2021–FY20262,86851 (1.8%)
Auditor's report, FY2016–FY20203,3193 (0.1%)
MD&A, FY2016–FY202012,9482,945 (22.7%)

The big annual book is almost always searchable. The auditor's report is almost always searchable. The short quarterly narrative from management is the one most often blind. My reading of why — and this is an interpretation, not something I measured — is that many MD&A filings are written as a letter to the exchange. The CENTEL filing is literally headed "LETTER OF CLARIFICATION for CENTEL's Operating Performance Results". Letters get printed, signed and scanned. Annual reports go through a designer and come out digital.

Why this matters: the MD&A is where management explains the numbers in its own words, quarter by quarter — it is the document you read to catch a story changing. It is also the document type most likely to be invisible to your search box. The blind spot is sitting on the most valuable pages, not the least.

4. The 10-second test: is this PDF blind?

You can check any filing yourself, with no software beyond your PDF reader. Do this before you trust a "no matches" result.

  1. Pick a word you can already see on the page. The company name in the header works. So does "Baht" or the year.
  2. Search for it. If a word that is plainly in front of you returns zero matches, the page has no text layer. Stop trusting every search you ran on that file.
  3. Try to highlight a sentence with your cursor. On a blind page, you will draw a box over the whole page instead of selecting words.
  4. Check the other language edition. Many Thai companies file both a Thai and an English MD&A. One of the pair is sometimes digital when the other is not — look for both on the company's filing page at the SEC.
  5. If both are blind, you have two honest options: read the pages with your eyes, or run OCR (optical character recognition) to rebuild a text layer. Free tools exist — ocrmypdf is a common one. Before you use it, read the next section.

Keep this habit: a search result of zero is a claim about the file, not about the company. Confirm the file can be searched before you let "nothing found" shape a decision.

5. The trap: OCR text is not the filing

Once you know a file is blind, the tempting fix is to OCR it and search the output. That fixes layer 3 and quietly creates a new problem: the OCR text is a machine's guess at what the page says, and it is wrong in ways that are hard to see.

Here is what a standard OCR engine produced from the exact paragraph I quoted at the top of this article, before any checking. I took this from our own pipeline's raw output:

What the page says (checked by eye)What raw OCR produced
"short term and long term loans totaling""short term and tong termลทรtotaling"
"support any operating expenses in the event""support any operating in the event"
"until the mid of next year""until the mic of next year"
(page 7) "provisions for long term employee benefits""provisions for fong term employee benefits"

Look at the first row. The word loans simply vanished, and three stray Thai characters appeared in its place — probably from a mark near the margin. Search the OCR output for "loans" and this paragraph is invisible again. You fixed the file and your search still misses the sentence.

Words are the gentle case. When OCR misreads a digit — a 3 for an 8, a dropped decimal point — it gives you a number that looks perfectly real and never existed. I wrote about exactly that failure in how AI fabricates financial numbers and the gates that stop it. OCR is the same risk, one step earlier in the chain.

What model validation teaches: a transformed dataset is not the source data. Any output built on OCR text needs a check against the original page, word by word, before it is allowed to count as a quote. Without that check, you have not recovered the filing — you have created a second document that merely resembles it.

6. What reading it all yourself actually costs

The honest alternative to search is to read. So how much reading is that? From the same files:

DocumentTypical length (median)Middle half of filesHow often filed
One Report (Form 56-1)255 pages205–317 pagesonce a year
MD&A5 pages3–8 pagesabout four a year (three quarters + year-end)

For one company over six years, that is about 6 × 255 + 24 × 5 ≈ 1,650 pages. For a portfolio of eight names, it is about 13,000 pages — before you open a single Opportunity Day presentation.

I am not going to put a reading speed on that, because yours is not mine. Pick your own pages per hour and do the division. Whatever number you get, two things are true:

  • The pages are public. Every file here is free to download from the SEC. Nobody is hiding anything from you.
  • The asymmetry is time. An institution with a team and a pipeline reads all 13,000 pages every quarter. An individual reads the latest quarter of the name that is currently worrying them. Same data, completely different view of it.

And that estimate assumes every page is searchable, which section 3 showed is not true — about a quarter of the older MD&A pile needs your eyes regardless of what tool you use.

Why this matters: "information asymmetry" in listed markets is rarely about secret information any more. It is about who can afford to read the public information at full depth. That is a cost problem, and cost problems can be engineered down.

7. How the Terminal handles the three layers

This is the part where I tell you what I built, so here is the plain version, including where it falls short.

Layer 1 — different words. The Terminal searches by meaning rather than by characters, so a question about liquidity reaches "credit facilities … yet still to be drawn down" without you guessing the phrase. The mechanics are in the semantic search article.

Layer 2 — same word, wrong meaning. Search results are ranked by how close the sentence is to your question, not by whether a word appears. "Repayment from loan to employees" does not sit close to "how much does this company borrow" in meaning, even though they share a word.

Layer 3 — no words. Scanned filings go through OCR with coordinates: every recognised word is placed back onto its position on the page image, and the placement is checked against the page before the text is used. Pages that are too skewed to place reliably are skipped rather than straightened and guessed.

The rule under all three: every result must point back to the exact PDF page or Opportunity Day timestamp it came from. If it can't be traced, we drop it. Not badge it, not show it with a warning — drop it.

Here is what that rule did to the CENTEL filing from the top of this article, as of 13 September 2026:

CENTEL MD&A Q2/2020 (scanned, 18 pages)Count
Statements extracted from the OCR'd pages59
Proven against the page (quote found, box drawn on the right page)48
Could not be proven — withheld from results11

One of the 48 is this line, which the system traced to page 6 — a page Ctrl+F could not read:

CENTEL · MD&A Q2/2020 · page 6: "Average RevPar decreased by 96.5% YoY to be at Baht 104.-, as a result of the Average Occupancy Rate (OCC) decreasing from 72.9% to 4.2% during Q2/2020; while Average Room Rate (ARR) decreased by 38.5% YoY to Baht 2,490.-"

— verbatim · company filings at the SEC

Notice the other side of that table: 11 statements from this one filing are not shown to anyone, because we could not stand behind them. That is a real gap, and I would rather you know about it than find out on your own. It is the price of the rule.

The Terminal is laser-focused on the MD&A, Risk, and Auditor sections — the parts of a filing where management and the auditor explain what the numbers mean — rather than trying to index every page of every document. It does not claim to be free of noise: meaning-based search always returns something that looks relevant, which is exactly why the page link sits next to every result.

Read it yourselfCtrl+FBoom Leverage Terminal
Different words, same meaningCatches it, slowlyMissesCatches it by meaning
Same word, wrong meaningCatches it, slowlyFloods youRanked by sentence meaning
Scanned, image-only PDFReadable by eyeBlindOCR'd, placed on the page, checked
Across years and companiesWeeks per portfolioOne file at a timeOne question, whole market
Can you check the source?YesYesYes — every result links to its page
Its own weaknessTimeEverything aboveWithholds what it cannot prove, so some statements are missing; semantic results still need your judgement

8. How far back you can see

The blind-file problem is worst in the older filings, so how deep your history goes decides how much of it the Terminal has already done for you.

PackageFiling historyOpportunity Day
DELTA4 yearsthe past 2 years
GAMMA6 yearsthe past 4 years
ALPHA10 years (FY2016–FY2026)every round in the archive · plus bulk export

ALPHA is the one that reaches back across the whole FY2016–FY2020 block counted in the table above — the era where roughly one MD&A in four is an image and a search box sees nothing. It is set up for teams and scoped in a conversation rather than bought off a card.

The big picture

Ctrl+F fails in three layers. It misses the euphemism. It floods you with the right word in the wrong meaning. And on 22.7% of Thai MD&A filings from FY2016–FY2020 — and still around 14% in recent years — it is searching a picture, finds nothing, and tells you so in exactly the same words it uses when the risk really is not there.

None of this is hidden information. All of it is a reading-cost problem. You can solve it by hand: run the 10-second test, read the blind pages with your eyes, and check any OCR against the page before you quote it. Or you can hand the pile to a system that does those steps and shows you the page for every answer.

Try it on a filing you already hold: ask the question you would have typed into Ctrl+F, in plain language, at Boom Leverage Terminal — then click through to the page and check it with your own eyes.

For a team or an institution that needs the full 10-year depth, bulk export or an API, talk to us at the Enterprise page or contact@boomleverage.com.

More reading: why meaning-based search beats matching characters — no executive ever types "the company is in trouble" · where in a 56-1 to open first — how funds actually read a One Report · what reading a management story across years looks like — PSL, FY2021–FY2026

Disclaimer: This article is produced for education and to explain how publicly disclosed filings can be read. It is not investment advice, it does not recommend any security, and it guarantees no outcome. Every company passage quoted here reproduces what the company wrote in a document filed through the SEC, with the page number given so you can check it. Quoting a company's disclosure is not a view on that company's financial health. File counts and percentages were measured from our archive on 13 September 2026 and will change as filings are added. Investment carries risk, and past results do not guarantee the future.

Read next