Working With AI Agents to Keep Fact-Dense Web Pages Fresh
How do you keep a fact-dense web page current when the facts change every week? Think of a Wikipedia article about something in the news. Wikipedia solves this with a swarm of editors. My pages need the same type of vigilance, and they have one editor: me.
I track drugs and clinical trials related to GLP-1s and obesity. Given the amount of hype and noise in this area, I think accessible, reliable, up-to-date information is critical. I recently created the GLP-1 Field Guides, a set of long-lived explainer pages about the GLP-1 drugs (the obesity and diabetes class behind Ozempic and Wegovy) and related classes. But I didn’t want to just create simple explainers, given that there are already good explanations of how GLP-1s work. I wanted the Field Guides to also serve as an always up-to-date reference for these medicines and related obesity drugs, drawing on my tracker’s data. From running GLP-1 Observer, my dashboard site, I know that new results, trial changes, and updates on drugs arrive every day in this area of medicine. With this in mind, how do I keep any page in the Field Guides from going stale overnight?
The simplest approach would be to just give an AI agent the page and ask it to check everything and update things that need updating. What’s wrong with this? Nothing, for pages that are relatively simple to check and have few facts. But the pages I am maintaining hold hundreds of checkable facts about drugs in clinical development, each of which expires at a different rate. An AI agent handed the page re-checks the sources that it sees and may report them as accurately quoted, but that doesn’t mean that the agent has checked what happened since those web pages were written. When the agent says that no updates are needed, I can’t tell if that means certain facts were unchanged or just unchecked.
The simple “check this page” approach is also inefficient. Changes for any given fact are infrequent, but with this many facts, pages change often. If less than 1 percent of facts move daily, you are spending a ton of tokens to find the ones that need editing. I had run enough early experiments to steer away from this approach in production.
Software engineers have a similar challenge: when you change part of a very complex system, how do you know that everything still works? Automated testing addresses this. Tests can be written for each part of the system and checked whenever anything changes. Here it is the software engineer making the change: the code moves while the world around it holds still. With reference pages, the prose holds still but the world moves underneath it. In both of these situations, having easy-to-run, repeatable checks spares a person or model from re-reading everything and hoping nothing gets missed.
What I wanted was a system with automated tests for prose. I am interested in this idea as a generalizable approach for fields where the pace of change is high. I think agents have only recently become capable enough to reliably help with this, if the system around them is carefully constructed. I’ve built a solution for the GLP-1 Field Guides. There are about 230 checkable passages across five pages, each registered with a “receipt”. A nightly deterministic (no LLM) job figures out which ones may have moved. Only those candidates go on to the next two stages. This has run nightly since June 2026. It had some flaws out of the gate. For example, its own verification was the main source of noise for the first couple of months, and real changes were easy to miss. I go into how I fixed that below.
The rest of this piece walks through what I built, and the diagram below shows the whole process. I think of it as an “LLM sandwich” with three main stages that run daily: deterministic checks first, an LLM in the middle for time-consuming verification and drafting edits, and human review last. At each stage the volume that needs attention gets filtered down, so that by the time it gets to human editorial review, I can handle the volume in well under an hour per day (often just minutes). In a week that I measured in August, 95 percent of what the checks raised got resolved without needing an edit drafted. The goal is to keep the editorial judgment that I find gratifying, and push the busywork of checking details to earlier stages so they can be dealt with by deterministic code or the agents.
The ledger: every claim gets a receipt
Before these stages can run, I need the prose in a form that code can check. I call each registered passage a claim, whether it is a sentence, phrase, or a short paragraph. The fact is what the claim asserts. Every passage that makes a checkable claim gets a “receipt”: an entry in a ledger (a structured text file) that records the passage verbatim, the exact location on the page, the fact it asserts, how volatile the fact is, the drugs and clinical trials that it names, any query of my GLP-1 tracker database that produced the fact, any web sources, and two dates: a last-checked date and a last-changed date.
For example, one claim on the amylin-combination page reads, as of this writing: “No GLP-1 + amylin combination is approved as of September 2026, but CagriSema (Novo Nordisk) is under FDA review after Novo filed its New Drug Application in December 2025.” Its receipt notes that it is checked against the web rather than my database, that it should be rechecked at least every 60 days, and that it was last checked September 11 and last changed August 9. When the FDA makes a decision, this claim goes out of date, and the receipt lets the system pick that up.
This type of ledger holding the receipts for a page is an order of magnitude larger than the page itself, in my case between eight and twenty-eight times the size depending on the page. But it’s a design decision that lets me build code to systematically check the pages. It also gives me two checks before a page can be published. Every passage in the ledger must appear verbatim on the page, and the recorded location for a claim must match where it shows up on the published page so a reader can click a passage and see its sources. If either of these checks fails, the page doesn’t build. Those checks prove every receipt is on the page; however, they don’t guarantee that every claim on the page has a receipt. Claims that do not can silently age. I guard against that gap with a whole-page pass that runs on a slower cadence, which I come back to later in this article.
Detection: three lanes, no agent
The first of the three stages is detection. There is no LLM involved in this one; it’s purely deterministic code. Three “lanes” check three kinds of claims. Lane 1 is for claims that can be checked against the database. It replays database queries and compares the current result against the previous night’s. This is snapshot testing for prose.
Lane 2 joins each claim’s drugs or trials to new events my GLP-1 tracker app has collected. That tracker pipeline collects trial events, approvals, and press every day, and an LLM labels them. That collection and labeling is outside the scope of this piece. If any of those new events involving a claim’s drugs or trials was labeled by the tracker pipeline as a development change, the claim needs verification. Routine press mentions get recorded but don’t trigger verification.
Lane 3 is for checking web-only claims. These carry a recheck date, depending on how volatile they are, after which they are considered expired and need re-verification.
This three-lane detection stage does two things: claims that are unchanged get their last-checked date updated to that day, and any claims that trigger a lane are written to a queue for the next stage. Nothing in this stage touches the page text; only the check dates have been updated at this point.
Verification, drafting, human review
The second stage is verification and drafting. Given that a fact may be used in several places on a page, claims are clustered by underlying fact, with these clusters stored in the ledger. Clusters are established during the whole-page pass I mentioned earlier, not by the nightly run. Each cluster has its own entry in the ledger, including the claims it covers, the fact involved, a note on how to check it, and what would make it go stale.
Each cluster’s fact goes through adversarial checking using a verifier agent and an adjudicator agent (a pair for each cluster). The verifier agent for a cluster receives the fact, its associated claims, and a ledger note about what would make this fact go stale. The verifier uses the tracker database and the web to check the fact and decide whether the fact is still correct, and if not, what it thinks it should be. The adjudicator agent treats the verifier’s verdict as a claim to re-verify. It independently re-checks the fact based on the queries and sources that the verifier used. In cases where the verifier agent said a fact has not changed, it goes beyond that to find new information to try to disprove the verifier’s conclusion. For example, in one run in late July, the verifier counted three approved generic versions of semaglutide in Canada, but the adjudicator ran its own check and found four based on regulators’ notices.
After the verifier and adjudicator reach their conclusion for a cluster and each of its claims, a final agent drafts any proposed edits for claims in clusters where the facts changed or where a specific occurrence’s wording was flagged even if the facts were judged still true. The result is a file of proposed edits, each with the exact current and replacement text. In the next stage, discussed below, I have a local review page that reads that file and shows me each edit with its sources, and controls to accept, reject, or modify.
I tried running the verifier and the adjudicator with a weaker model. The weaker agents confirmed stale claims as still true three runs in a row. Once I added a stronger model to spot-check their work, the end-to-end process cost more than just using the strong model throughout. Even with this updated process, in August a day’s automated steps ran between about 500K and 2.5M tokens, depending on how much was flagged, which is the equivalent of eight to forty copies of The Great Gatsby.
The third stage is human editorial review. This step is necessary. For one, when I have rejected an edit, the reason for rejection doesn’t carry forward to future runs. For example, the verifier proposed an edit that I rejected on August 18. Two weeks later, a lighter version of the same edit was proposed as a new edit. I accepted it, then reverted it the same day after finding my earlier ruling in the record. Although the reasons behind decisions are recorded, feeding them forward is harder than it sounds. A note that says “previously settled” is sometimes interpreted by a model as an instruction to stop checking. The connection between present decisions and future runs is me at the review step.
Another reason that this human review step exists is for a view of the whole page. Using the claim and its fact as the unit of verification makes the system possible. But length is a property of the page. Before I added checks to the process, pages grew between 7 and 15 percent in length between May and August 2026 because edits tended to be additive, rarely neutral or subtractive. One July spot-check drafted seven edits that added 646 characters. By then, I was aware of the issue and minimal substitutions during my editorial pass kept all the fixes at zero net gain in length. I shipped fixes for this in mid-August: a rule that corrections prefer replacing text instead of adding to it, a per-page word budget where going over it triggers a warning, and word count deltas flagged for me on the review surface. This has stabilized the length of the pages.
Coherence is another page property. Every claim on a page can be true while the page contradicts itself or doesn’t flow well between paragraphs. A separate pass reads each whole page at a lower frequency, checking style, dates, and references together, and once a month a smaller pass moves every “as of” date forward.
In these cases, the process was blind to something I picked up on, like an earlier decision or the page growing in length. Keeping myself in the process guards against gaps like these.
Agents don’t update the page on their own. Each change goes through my review, and a script applies it if I accept. Agents may be capable enough to publish without human review at some point, but it’s not there yet in a field like clinical drug development.
When verification becomes noise
The second part of this piece is about the experience of running the process described above, and the ways in which I got it wrong. My goal is to automate as much of the process as I can and to enter it only where I am needed. The issues below caused me to spend more time than needed tending to the automations and their output.
In the hospital, “alarm fatigue” occurs in settings like the emergency department or ICU when monitors for heart rate, oxygen saturation, and other vital signs sound often. Clinicians learn to tune them out, and a critical alert can be missed among the routine ones. In my case, every query that an agent wrote while verifying a claim was added to the claim’s receipt, and the nightly check replayed all of them. For example, the claim at the top of the verification queue, the count of drugs in development that sits on every page, was the one I had checked most often. Each check stored more queries on its receipt, and each stored query was another chance for the next night’s replay to flag it. By July 24, 1,597 queries replayed nightly, and one claim carried 87 of them. More than half of the “changed” triggers that week, 228 of 414, fired on queries whose row count had not moved at all. Over four nights in late July, the nightly check produced 102 to 159 flags per night, against 224 claims on the whole site, and I was at risk of accepting agents’ edits without looking as carefully as I should have. The hospital equivalent would be a patient wearing 20 heart monitors, each with its own alarm, instead of one.
I found the problem on July 24 and shipped the changes right away. I capped the number of queries for each claim, I de-duplicated queries, and for count claims (for example, how many in-development drugs of this class there are), I kept one query, which flags only when the count crosses the number on the page. After these changes, nightly queries went from 1,597 to 289 and daily flags went from roughly 100 down to 17. It’s far easier now to tend to the process and give my attention to the changes that need it.
When I analyzed the first 52 nights on August 4, the numbers bore out the problem: there had been 1,228 flags, more than 400 of them in the four peak nights alone, and 63 applied edits. About one flag in twenty became an edit. More than three quarters of claims had never changed. Of 646 re-verifications, 2 concluded that something had changed. At that volume, the review step had not been sustainable at the quality that I am aiming for.
The bug I fixed five times
Of the 81 fixes I shipped in the first two months, just over half were fixes for symptoms of the same underlying bug. A query that failed was recorded as a result that had not changed, so the claim read as checked. A flag I had not dealt with on the day it was raised would sometimes clear itself on the next run. I made fixes in five different places in the code before I realized that the same underlying assumption was responsible: an error during checking, or no check at all, was recorded the same way as something that had been checked. The ledger didn’t have a state that could distinguish “could not check” from “checked, unchanged,” so every failure was recorded as the latter. The fix was to add that extra state. After addressing the underlying issue, failed checks show up as failed and flagged claims stay flagged until properly cleared. One centralized function answers whether a claim is verified as of today.
Hardening the process
Running this over time and reading the output taught me other lessons. Here are a few.
An alarm that goes off every night ceases to be an effective alarm. This is similar to the “alarm fatigue” issue I wrote about above that was due to queries accumulating over time. There the problem was the volume of checks. In this case, it was the quality of checks – whether what they were watching was specific enough. Any change in a replayed query’s result flagged as “high severity”, whatever the change was, and the nightly summary that tells me whether to review that day’s findings said yes seven nights out of seven. The fix was to decide in advance what would make each claim wrong, and to flag as high only on that.
For example, one claim stated that retatrutide was in Phase 3 for obesity. Its query returned that drug’s trial records, so any update to any of them, like a new completion date or an enrollment change, flagged the claim as high. The rule now watches only the phase and status fields, and changes to anything else in those records don’t trigger an alarm. Be clear on what the check is checking.
Once high meant something, I could move my editorial review to every 48 hours, with a same-day review only when a high alarm fires. Detection is cheap and still runs every night. This is one of the benefits of having deterministic checks be the first step in the process.
An early version of the system parsed each stored query in advance and decided whether it was safe to run. Queries it judged not runnable were dropped from the nightly run without raising a flag, so a claim could lose its watch and become a blind spot. I removed the pre-run checker and let the nightly run try each query. A query that fails now raises a flag instead of disappearing. Sometimes it’s better just to fail loudly and let the failure dictate the fix.
The page prose is not the only thing that ages. The notes that I wrote into the ledger also go out of date and need their own expiration dates. One note I had written on a cluster asserted a point that the page had moved past two weeks earlier, but its presence had the agents every day re-arguing a question that was already settled. With undated notes, I ended up wasting time and tokens, and risked having the agents suggest wrong edits. Now each note carries the date the judgment was asserted, and the system warns when a note’s date is older than the last change to its claims.
Who watches the watcher?
For almost two months, the lane that reads press releases had not been returning recent news stories because of a query bug. Nothing in the system was designed to alert me when silence was a sign of trouble. I added the checks below to keep this kind of silence from going unnoticed again.
A quiet night is ambiguous: it could mean that everything checked out fine or it could mean that the detector is broken. To tell the difference, after the real run each night, I take four probes with known answers and run them through the detection code. Three should raise a flag and one, a stable, correct fact, should not. If any of these probes gives a wrong answer, a warning goes out. An all-clear means the detector was tested that night and passed. Similarly, a separate check runs shortly after the detection lanes run, reads the last-checked time of every page, and alerts if the oldest one is more than twelve hours old. The oldest is the one to check; the newest ones look fine when a run dies halfway through. These checks are automated, and they let me trust that the system is still running even on a quiet night.
Those checks prove the detector is running. Another check looks at whether it’s picking up major changes. Each week I independently curate the material stories in obesity drug development for another part of the site. This recall check takes every story that names a drug or trial one of the pages covers and asks whether the detector raised a related flag. The last time I ran it, on September 14, 56 of 58 such stories had one. Any story that doesn’t have a related flag in the nightly checks becomes a question for me to investigate: did the page need to change? In this latest case, I checked the two that didn’t have associated flags. Neither page needed to change.
Takeaways
It would have been much harder to build this system a year ago, and I’m not sure it would have run reliably even a few months ago. Experimenting with and running the agents on real work every day is how I learned which tasks they are now able to do reliably. Some lessons, with the caveat that the system is still changing:
- The claim is the useful unit of checking, not the full page. A page with many facts is difficult to check as a page. Registering each claim with the fact that it asserts provides a structure that makes checking feasible for an agent system.
- Start with deterministic checks and reserve the model for the steps that need judgment. Most days, the cheap checks rule out most of the work before agents are called.
- Keep a person as the last step for now. Being in the loop gives me a chance to increase the quality of what my users see and to learn where the process needs improvement. I’ll know when the agents get good enough that I’m not adding value, but we’re not there yet.
- Don’t let verification add to the next detection run. This is one way to get into a positive feedback loop, with an overwhelming amount of noise before long.
- Give every state its own representation. Distinct states like “could not check” and “checked, unchanged” need to look different to the system.
- Reserve high-severity flags for changes that would make a claim wrong. Alarms that go off every night for minor changes start to be ignored over time and quickly become useless.
- Silence doesn’t mean all is well with the system. Build checks that check the detector itself.
- Check the measuring instrument before trusting its stats. The detector was blind to press articles for two months while I drew conclusions from its numbers.
- Some properties like length, consistency, and coherence belong to the whole page. Checking and fixing individual claims can’t address these properties; whole page checking helps keep quality high.
I expect agents to keep getting better quickly, but also recognize where we are on that arc and feel a sense of responsibility to readers to keep the quality high and have these pages be the best resource on the topic out there. The bar for agent-assisted work should be higher than what a person could produce alone. Coming back to the Wikipedia example, most of us as creators don’t have access to a swarm of editors, but we do now have a team of research, checking, and drafting assistants available to us in the form of agents that can help keep pages current on fast-moving topics, if the system around them is built with care.