Method
Everything on this site is produced by a documented, repeatable process. The numbers below are read from the running configuration when this page is built, so they cannot drift out of step with what the system actually does.
What is collected
Coverage is pulled from two sources every day: Google News search feeds across 14 national editions (US, UK, Australia, Canada, India, Ireland, South Africa, Singapore, Nigeria, New Zealand, Philippines, Kenya, Pakistan, Malaysia), and direct RSS feeds published by individual outlets. Direct feeds matter because they supply the real article URL, which is what makes body text and archiving possible.
15 standing searches run each day:
| Watchlist | What it looks for |
|---|---|
anonymous_sourcing | Reporting resting on sources who are not named. The single highest-value technique list: it applies to every beat, every outlet and every political direction, and techniques.py already extracts it per article. Whether anonymity was justified is not ours to say; whether it was used, and whether a reason was given, is observable. |
unnamed_authority | Appeals to authority with no identifiable authority attached. Catches "experts say" and "critics say" equally, which is the point: both are the same construction pointed in opposite directions. |
projection_stated_as_outcome | Forecasts, models and plans reported in language that reads as settled. Applies to budget projections, weather, epidemiology, markets and policy alike. |
quantified_claims | Numbers in headlines, where the comparison base is the thing most often missing. "Doubled" from what, "record" since when, "times more likely" than whom. |
health_and_medical_claims | Where a reader has the highest stake and the least ability to check: single studies reported as findings, relative risk without absolute risk, recalls, approvals, trial results. |
economy_and_cost_of_living | Indicators that move policy and household decisions, and that are routinely reported without the revision history or the margin. |
polling_and_elections | Poll coverage, where the margin of error, the sample and the question wording are the story and are usually absent. Written to catch any lead, in any direction. |
science_and_environment | Research and environmental reporting, including both alarm and dismissal framings. Catches "scientists warn" and "scientists dispute" together. |
local_government_and_budgets | The layer of government with the most direct effect on a reader and the least coverage anywhere. Almost never syndicated, so it is also a useful contrast case against wire-driven beats. |
criminal_charges_framing | Court and criminal coverage. Retained, but no longer the majority of the intake. |
relational_descriptor_leads | Headlines leading with a family or occupational role rather than the conduct at issue. |
passive_voice_incidents | Passive construction that removes the actor from the sentence. |
immigration_status_framing | Coverage where immigration status is either foregrounded or omitted. Catch both choices. |
corrections_and_retractions | Outlets acknowledging error. This is the positive-signal feed and it feeds the "where they got it right" record: an outlet that corrects itself belongs in the ledger as prominently as one that does not. |
official_statements_vs_coverage | Cases where a primary document exists, so coverage can be compared against the source rather than against other coverage. |
These are written to catch patterns regardless of who the subject is. A search that could only ever surface one political direction would make this site the thing it criticises.
How coverage is grouped
Articles covering the same event are grouped so their headlines can be compared. Grouping requires a shared proper noun — a name, a place, an institution — appearing in no more than 1% of the day's articles.
This matters more than it sounds. "Man charged over fraud in London" and "Man charged over fraud in Dublin" share four words of five and are unrelated crimes. Word similarity alone groups them together. Requiring a shared rare proper noun does not.
| Rule | Value |
|---|---|
| Minimum distinct outlets | 4 |
| Maximum outlets before a group is discarded as a topic bucket | 15 |
| Maximum articles per outlet | 2.5 |
| Groups analysed per day | 36 |
How outlets are counted
Every entry gives two numbers: how many mastheads published a story, and how many independent sources those mastheads represent. Outlets collapse into one source when either is true:
- Identical text. Word-for-word matching headlines indicate one origin redistributed — a wire pickup or a chain pushing copy to every station it owns. Detected from the text itself, so it needs no maintained list and catches outlets we have never heard of.
- Shared ownership. 20 corporate groups covering 131 outlets are recorded. Anything unlisted is treated as independently owned, which is the generous assumption — gaps make our counts too high, never too low.
Outlets are also tiered by type:
| Tier | Outlets classified |
|---|---|
wire | 10 |
public | 17 |
independent | 62 |
state | 18 |
investigative | 5 |
partisan | 12 |
Unclassified outlets are shown as unclassified rather than assumed trustworthy.
What the analysis does and does not do
Language models are used for extraction, never for judgement. They are asked what a headline states, which facts appear in the body but not the headline, and which outlets assert a thing. They are never asked whether a headline is biased, what a writer intended, or whether a claim is true.
Extraction is checkable in seconds by anyone who clicks the link. Judgement is checkable by nobody. Every entry here is built only from the first kind.
Language audit
Where body text is available, articles are audited for observable constructions: assertions with no attribution, anonymous sourcing, attribution verbs, descriptive modifiers, causal claims without supporting evidence, and whether the other party's response appears. Findings quote the exact text. Articles with no findings are published alongside those with findings.
Emotional register
Headlines are scored for negative and positive language, unresolved threat, and whether an outcome is reported. Comparisons are always against every other outlet on the same days, so a genuinely grim week raises all outlets and cancels out.
A high absolute rate is not evidence of anything. Journalists cover aberrations and aberrations skew bad. What is informative is one outlet running well above the others on identical events, or consistently raising threats it never resolves.
Post-publication changes
Every headline observed is logged with a timestamp. When the same URL later shows different text, both versions are kept with a word-level diff.
Timestamps mark when this system looked, not when an edit was made. A change is reported as having occurred within a window between two observations. Entries say "changed within a 5-hour window", never "changed 5 hours after publication", because the second is a claim the data does not support.
This project’s own record
Every headline seen here is logged against its URL with the time it was seen. When one changes, both versions are already on record, together with the window in which the change happened. That log is this project’s own: it is held locally, and it is hash-chained into the seal, so any alteration to an earlier entry breaks every seal published since.
It has no page of its own to link to. Where a verification rests on it, the source is cited as sellmefear headline capture record and carries no external link — not because a link was omitted, but because there is no outside document. The record is ours, and the seal is what makes it checkable.
Every change it holds is listed on the changes page.
Verification
Claims are sorted into two kinds.
- Claims about what an outlet published are self-verifying — the outlet is the primary source for its own words, and the archived snapshot is the evidence.
- Claims about what happened require a non-media primary source: court records, agency filings, official statistics, corporate registries. A news article is never cited to establish a fact about the world.
Where a fact could not be confirmed against a primary source, it is published as unconfirmed rather than repeated as true.
Archiving
Every source is captured three ways when first seen, not when published here: a Wayback Machine snapshot, a screenshot of the page as it appeared, and the stored HTML. References are listed in that order because originals move and disappear. Where a source exists only as a live link, it is marked as having no durable copy and is not cited.
Integrity of the record
Every record is hash-chained to the one before it, so altering or deleting any past entry breaks the chain. A fingerprint of the whole ledger is published to a public repository on a schedule, which timestamps it through a third party rather than through us.
This exists because a record that asks to be trusted is the same instrument as the one this site was built to answer. Nobody should have to take our word for what we recorded.
What leads the front page
The edit at the top of the front page is the one where the most of the headline changed, among changes observed in the last seven days that have not already appeared there.
That is a measurement rather than an opinion. For every change, the proportion of the original wording that survived is computed and published beside it; the edit with least surviving wording leads. Anyone can count the words and check the ordering, in the same way anyone can check a timestamp.
Three things this rule deliberately does not consider: the subject of the story, the outlet that published it, and whether the edit seems important. The last is the one that matters most. An edit that changes half a headline may be trivial and one that changes a single word may be serious, and deciding which is which would be a judgement a reader has to take on trust. This site does not ask for that.
The seven-day window keeps the page current, since the largest change ever recorded would otherwise sit here permanently. An edit does not lead twice, so an outlet serving two versions of one headline cannot hold the position. Everything remains on the changes page on equal terms.
Selection, and its limits
Choosing which stories to examine is an editorial act, and no methodology removes that. Two things constrain it: the watchlists are fixed in advance and published above, and every candidate the system surfaced that was not published is listed with the reason.
Known limitations
- Body text can be retrieved for roughly half of articles. The rest are assessed on headline alone, and each entry states which.
- The record has no history before this system started running and cannot be backfilled.
- Coverage is English-language.
- Outlets are classified local, national, or international by circulation; anything unclassified defaults to local.
- Grouping errors are possible. Every entry lists its sources so a misgrouping is visible rather than hidden.
Generated 27 August 2026 from the running configuration.