// 2026-10-08
The filter was fine. The adapter had been lying for months.
A scheduled job started reporting nothing. The obvious explanation was wrong, and finding that out took one measurement instead of one guess.
I run a small system that sweeps job boards and discussion threads overnight, scores what it finds with an LLM, and mails me one ranked digest. For six days it reported the same thing: zero.
The numbers it printed were scanned=213, filtered=1. Two hundred and thirteen items in, one survived the keyword filter. The obvious reading is that the filter is broken.
The obvious reading was wrong
Before changing anything I measured it. I pulled the live source directly, ran the same filter over it outside the pipeline, and counted:
102 of 200 passed. The filter was working exactly as intended.
That one number killed the plan I had been about to execute, which was to go and loosen filter rules that were not the problem. Loosening them would have let through more noise, the digest would have got worse, and I would have concluded the scoring was at fault next.
Where it actually was
The biggest source is a monthly discussion thread with a few hundred comments. The adapter asked the API for its first page and read the results. It never looked at the field next to them saying how many pages there were.
nbPages: 2
Every run since that adapter was written had been reading exactly half of the largest source in the system, and nothing anywhere said so. No error, no warning, no gap in the output — just a number that was quietly smaller than it should have been.
Why it stayed hidden
Because the only thing the system reported was a total. A total cannot distinguish a quiet day from a dead source, and it cannot distinguish a source returning everything from a source returning half.
The fix was two lines of pagination. The change worth keeping was to the reporting: the digest now prints a per-source breakdown with the empty ones in bold, and says "no sources reached" rather than "0 new" when nothing was reachable at all. A source that returns nothing is now visibly a source that returned nothing, rather than an absence inside a sum.
The habit
I have started treating "the obvious cause" as a hypothesis with a cost attached. Measuring the filter took about ten minutes. Rewriting it would have taken an afternoon and made the system worse.
The question that got me there was not clever. It was just: *before I change this, what does it actually do right now?*