sellmefear

Method

Everything on this site is produced by a documented, repeatable process. The numbers below are read from the running configuration when this page is built, so they cannot drift out of step with what the system actually does.

What is collected

Coverage is pulled from two sources every day: Google News search feeds across 14 national editions (US, UK, Australia, Canada, India, Ireland, South Africa, Singapore, Nigeria, New Zealand, Philippines, Kenya, Pakistan, Malaysia), and direct RSS feeds published by individual outlets. Direct feeds matter because they supply the real article URL, which is what makes body text and archiving possible.

15 standing searches run each day:

WatchlistWhat it looks for
anonymous_sourcingReporting resting on sources who are not named. The single highest-value technique list: it applies to every beat, every outlet and every political direction, and techniques.py already extracts it per article. Whether anonymity was justified is not ours to say; whether it was used, and whether a reason was given, is observable.
unnamed_authorityAppeals to authority with no identifiable authority attached. Catches "experts say" and "critics say" equally, which is the point: both are the same construction pointed in opposite directions.
projection_stated_as_outcomeForecasts, models and plans reported in language that reads as settled. Applies to budget projections, weather, epidemiology, markets and policy alike.
quantified_claimsNumbers in headlines, where the comparison base is the thing most often missing. "Doubled" from what, "record" since when, "times more likely" than whom.
health_and_medical_claimsWhere a reader has the highest stake and the least ability to check: single studies reported as findings, relative risk without absolute risk, recalls, approvals, trial results.
economy_and_cost_of_livingIndicators that move policy and household decisions, and that are routinely reported without the revision history or the margin.
polling_and_electionsPoll coverage, where the margin of error, the sample and the question wording are the story and are usually absent. Written to catch any lead, in any direction.
science_and_environmentResearch and environmental reporting, including both alarm and dismissal framings. Catches "scientists warn" and "scientists dispute" together.
local_government_and_budgetsThe layer of government with the most direct effect on a reader and the least coverage anywhere. Almost never syndicated, so it is also a useful contrast case against wire-driven beats.
criminal_charges_framingCourt and criminal coverage. Retained, but no longer the majority of the intake.
relational_descriptor_leadsHeadlines leading with a family or occupational role rather than the conduct at issue.
passive_voice_incidentsPassive construction that removes the actor from the sentence.
immigration_status_framingCoverage where immigration status is either foregrounded or omitted. Catch both choices.
corrections_and_retractionsOutlets acknowledging error. This is the positive-signal feed and it feeds the "where they got it right" record: an outlet that corrects itself belongs in the ledger as prominently as one that does not.
official_statements_vs_coverageCases where a primary document exists, so coverage can be compared against the source rather than against other coverage.

These are written to catch patterns regardless of who the subject is. A search that could only ever surface one political direction would make this site the thing it criticises.

How coverage is grouped

Articles covering the same event are grouped so their headlines can be compared. Grouping requires a shared proper noun — a name, a place, an institution — appearing in no more than 1% of the day's articles.

This matters more than it sounds. "Man charged over fraud in London" and "Man charged over fraud in Dublin" share four words of five and are unrelated crimes. Word similarity alone groups them together. Requiring a shared rare proper noun does not.

RuleValue
Minimum distinct outlets4
Maximum outlets before a group is discarded as a topic bucket 15
Maximum articles per outlet 2.5
Groups analysed per day 36

How outlets are counted

Every entry gives two numbers: how many mastheads published a story, and how many independent sources those mastheads represent. Outlets collapse into one source when either is true:

Outlets are also tiered by type:

TierOutlets classified
wire10
public17
independent62
state18
investigative5
partisan12

Unclassified outlets are shown as unclassified rather than assumed trustworthy.

What the analysis does and does not do

Language models are used for extraction, never for judgement. They are asked what a headline states, which facts appear in the body but not the headline, and which outlets assert a thing. They are never asked whether a headline is biased, what a writer intended, or whether a claim is true.

Extraction is checkable in seconds by anyone who clicks the link. Judgement is checkable by nobody. Every entry here is built only from the first kind.

Language audit

Where body text is available, articles are audited for observable constructions: assertions with no attribution, anonymous sourcing, attribution verbs, descriptive modifiers, causal claims without supporting evidence, and whether the other party's response appears. Findings quote the exact text. Articles with no findings are published alongside those with findings.

Emotional register

Headlines are scored for negative and positive language, unresolved threat, and whether an outcome is reported. Comparisons are always against every other outlet on the same days, so a genuinely grim week raises all outlets and cancels out.

A high absolute rate is not evidence of anything. Journalists cover aberrations and aberrations skew bad. What is informative is one outlet running well above the others on identical events, or consistently raising threats it never resolves.

Post-publication changes

Every headline observed is logged with a timestamp. When the same URL later shows different text, both versions are kept with a word-level diff.

Timestamps mark when this system looked, not when an edit was made. A change is reported as having occurred within a window between two observations. Entries say "changed within a 5-hour window", never "changed 5 hours after publication", because the second is a claim the data does not support.

This project’s own record

Every headline seen here is logged against its URL with the time it was seen. When one changes, both versions are already on record, together with the window in which the change happened. That log is this project’s own: it is held locally, and it is hash-chained into the seal, so any alteration to an earlier entry breaks every seal published since.

It has no page of its own to link to. Where a verification rests on it, the source is cited as sellmefear headline capture record and carries no external link — not because a link was omitted, but because there is no outside document. The record is ours, and the seal is what makes it checkable.

Every change it holds is listed on the changes page.

Verification

Claims are sorted into two kinds.

Where a fact could not be confirmed against a primary source, it is published as unconfirmed rather than repeated as true.

Archiving

Every source is captured three ways when first seen, not when published here: a Wayback Machine snapshot, a screenshot of the page as it appeared, and the stored HTML. References are listed in that order because originals move and disappear. Where a source exists only as a live link, it is marked as having no durable copy and is not cited.

Integrity of the record

Every record is hash-chained to the one before it, so altering or deleting any past entry breaks the chain. A fingerprint of the whole ledger is published to a public repository on a schedule, which timestamps it through a third party rather than through us.

This exists because a record that asks to be trusted is the same instrument as the one this site was built to answer. Nobody should have to take our word for what we recorded.

What leads the front page

The edit at the top of the front page is the one where the most of the headline changed, among changes observed in the last seven days that have not already appeared there.

That is a measurement rather than an opinion. For every change, the proportion of the original wording that survived is computed and published beside it; the edit with least surviving wording leads. Anyone can count the words and check the ordering, in the same way anyone can check a timestamp.

Three things this rule deliberately does not consider: the subject of the story, the outlet that published it, and whether the edit seems important. The last is the one that matters most. An edit that changes half a headline may be trivial and one that changes a single word may be serious, and deciding which is which would be a judgement a reader has to take on trust. This site does not ask for that.

The seven-day window keeps the page current, since the largest change ever recorded would otherwise sit here permanently. An edit does not lead twice, so an outlet serving two versions of one headline cannot hold the position. Everything remains on the changes page on equal terms.

Selection, and its limits

Choosing which stories to examine is an editorial act, and no methodology removes that. Two things constrain it: the watchlists are fixed in advance and published above, and every candidate the system surfaced that was not published is listed with the reason.

Known limitations

Generated 27 August 2026 from the running configuration.