/methodology · how the data is verified

How football facts become verifiable.

Verifiability is the goal and the standard we measure against—not a blanket claim about today’s database. foot.io records sources and confidence where the ingestion path supplies them, compares independent sources for supported datasets, and publishes the gaps that remain.

This page separates the implemented controls from the work in progress: source attribution, cross-checking, canonical identities, validation and corrections.

Provenance, measured honestly

01 · where each value came from—and where evidence is missing

The canonical schema supports row-level source columns, external identifiers and a fact-evidence ledger. Importers are expected to write source and retrieval metadata rather than leave an untraceable value. That contract is not yet satisfied everywhere: evidence records are concentrated on matches, source URLs are sparse, and several entity classes do not yet have fact-level evidence.

Every new entity-fact assertion must resolve to the registered source catalogue, name an existing canonical subject and object, and carry a stable source record identifier or retrievable URL. When an importer has an audit lease, the assertion also records the exact import run; retrieval time and a SHA-256 content digest are supported for reproducible snapshots. Query entity_fact_provenance for the evidence and its honest status. Historical rows remain labelled source_only or import_run_only where stronger evidence does not exist—we do not manufacture provenance retroactively. Dangling legacy assertions leave public results only after their complete pre-images are retained in an internal integrity quarantine.

Completeness needs a denominator

02 · missing seasons must exist in the plan before they can be measured

Counting only seasons already in the database makes absence invisible. The expected-season catalogue therefore gives every public competition an explicit research state. Source-backed targets expand from their first expected season through the last completed season reviewed; non-held years remain visible as sourced exceptions. The API then distinguishes a missing season row, an empty season, a season with no scores, duplicate season identities and a season with played matches.

An unresearched competition is never included in a flattering percentage, a provisional target can expose gaps but cannot claim completion, and only a verified target whose review is current and whose entire expected range is present can set verified_complete=true. Canonical identity, provisional lifecycle research and a verified coverage contract are separate gates. A provisional range needs source-backed inception plus an explicit end date or stored-season evidence; absence of a dissolution claim is never interpreted as current activity. Wikipedia can support only the provisional gate, only through an exact canonical identity, compatible competition infobox and immutable article revision. Automated refusals remain queryable with their method, reason, retry date and next evidence task; they are not silently rerun or counted as successful research.

The backlog is prioritised without weakening that rule. Unresearched competitions are ranked by indexed-match impact, current relevance and tier, while every known problem year is routed to a named ingestion or integrity-repair lane. A candidate inception year remains a research lead until a retrievable source supports the lifecycle contract.

The same rule now applies inside stored seasons. Every public competition-season has a fixture-contract state. A cited expected total is compared with live row count, an orientation-independent distinct-fixture count, scored matches and scoreless fixtures; every season without that source-backed denominator is explicitly unresearched. A full API schedule response can create or refresh the contract automatically, while a partial paginated response cannot.

Completeness metadata is public, but raw historical entitlement is separate. Postgres RLS restricts canonical season-, match-, event-, lineup- and match-stat rows to current seasons for the public preview and free keys; historically entitled keys and settled paid MCP calls can read the full depth. Keyed MCP queries traverse the same Postgres policy instead of bypassing it with an internal service role.

Classification is a fact contract

03 · unknown stays unknown until evidence supports it

A league-looking name is not evidence that a competition is a senior domestic first tier. Format, gender, age group, geographic scope, governing country, pyramid tier and lineage are recorded as separate source observations. The compatibility columns remain useful for queries, but a populated value without retrievable source evidence is explicitly unverified.

The public classification queue distinguishes a missing required value from missing evidence and from a genuine source conflict. Conflicts fail closed: all candidates and their sources remain queryable, while no source silently overwrites the canonical value. Importers may promote only fields explicitly supplied by the source; derived defaults and name-pattern guesses are not accepted as proof.

Raw CSV and bulk rows are isolated in a private staging schema. The legacy name-only promoter is retired. A staged match now needs an active registered source, stable source match and record IDs, retrieval time, SHA-256 content digest, one resolved competition-season and two compatible existing teams before an atomic database function can promote it. A staged player must resolve to an existing canonical player; ambiguous names remain quarantined rather than creating or fusing an identity. Verified publisher IDs are stored in a uniqueness-checked source_entity_identities registry before legacy JSON stamps are considered.

Cross-source corroboration

04 · we don't trust a single source

No single provider is authoritative for the whole sport, so we don't treat any one as gospel. Where independent sources can be compared, we compare them — and we label the result honestly with one of three tiers.

The entity-fact ledger exposes single_source, multi_source_unverified, corroborated and conflict states. Corroborated requires agreement across at least two explicitly assessed publisher-independence groups: FBref plus Stathead, for example, remains one Sports Reference group. A scalar conflict—two current values for attendance, result, venue or another one-value predicate—is omitted from the resolved endpoint until reviewed; every candidate remains queryable with its sources. Multi-valued predicates such as a starting lineup resolve every membership independently.

Historical match assertions also have an explicit partial state. If a publisher supplies a stable round, teams and score but no calendar date, foot.io keeps the canonical date null and stores the round as the occurrence discriminator instead of inventing a date. Partial assertions are machine-readable evidence but never count as date-verified corroboration. Conversely, a fully identified publisher result can complete an empty canonical score only through a journalled, fail-closed promotion; populated results are never overwritten.

Second-publisher acquisition retains two denominators: every mapped single-source match still owed and the subset actionable against the chosen publisher now. Unsupported, missing-ID and no-exact-match attempts receive a retry date; they are deferred from scheduling, never misreported as verified or removed from the full gap.

Match-scoped facts also fail closed across table boundaries: an event, lineup, player-stat or shot row cannot name a team that did not play the parent match, and a match’s sides cannot be edited if that would orphan an existing fact. Contradictory legacy rows are removed from canonical results only after their full pre-images are written to an internal integrity quarantine for source-backed review and restoration.

Match events are unified separately from their source observations. Two providers describing the same player, minute and event type produce one canonical event plus two immutable evidence payloads—not two goals in the API. The match_event_verification view reports its sources and labels optional-field disagreements as conflict; disputed fields remain empty until review.

Assists are definition-bound: publishers do not all mean the same thing by “assist”, while a potential assist includes the action preceding a teammate’s shot whether or not the shot is converted. We store the definition and version beside every assertion. Rule- and model-derived candidates retain their input evidence but cannot masquerade as publisher reports. If two publishers name different assisters, the event fails closed as a conflict.

Human annotation is method provenance, not extra publishers: two raters inside one managed annotation batch can establish inter-rater agreement and support adjudication, but they do not count as two independent sources. Batches record non-personal QA and labour metadata—method, rater count, worker geography, engagement model and whether compensation is disclosed—without creating a worker identity registry.

Live measurement: exact publisher-aware totals are served by fact_tier_summary. Earlier source-label percentages were withdrawn because aliases, metadata keys, FBref plus Stathead, and derived rollups could overstate independence. These remain row-level tiers, not verification of every individual field.

Magnitude, measured on 6 August 2026: corroboration is early, and we publish how early. 2 entity facts currently carry the corroborated tier against 3,591,643 single-source; 81 canonical match events are corroborated against 1,745,004 single-source; 21 of 3,903 competition coverage targets are verified, the rest provisional or not yet researched. The product today is the label, not implied mass verification — every fact states honestly whether it is corroborated, single-source or in conflict, so you can filter to the tier your use case needs. The corroborated share grows as second-publisher acquisition works through the queue; the labels will report that growth rather than assert it in advance.

Corroborated
Two or more independent sources agree under the tier view’s comparison rules. This is stronger evidence, but still requires checking the cited sources for high-stakes use.
≥ 2 publisher groups agree
Single-source
Carried by one source and not yet cross-checked. Usually correct, but not independently confirmed — and we say so rather than imply more certainty than we have.
1 source, unconfirmed
Flagged
Sources disagree, or a value failed a plausibility check. We keep both values with their sources and queue it for review — we never silently pick a winner.
conflict / needs review

One clean identity per entity

05 · entity resolution & de-duplication

The hardest part of unifying football data isn't the stats — it's deciding that "Man Utd", "Manchester United FC" and a Wikidata ID all refer to the same club, while keeping the men's and women's sides distinct. foot.io resolves these through an alias table and dedicated resolvers before anything is written, so each real-world club, player, competition and stadium maps to exactly one canonical record.

Gender-aware
Never conflated
Matching is gender-aware end to end — a women's competition routes to the women's side of a club, never the men's. Postgres rejects a match whose explicitly gendered team conflicts with its competition, and rejects parent edits that would create the same mismatch.
Era-aware
Same name, different club
Clubs fold, merge, and rename. Two teams that shared a name across different eras are matched by date, not by string — so a 1900 result isn't accidentally welded onto a modern namesake.
Reversible
We un-fuse mistakes
When we discover that two identities were wrongly merged, we split them back apart and reattribute their history to the correct entity — rather than papering over the join.
Staged
Validated before promotion
The target contract is staging-first promotion after validation. Some legacy ingestion paths still write directly to canonical tables; moving those paths behind one auditable promotion gate is active remediation work.
Run-isolated
Incomplete never means green
Managed ingestion runs have a database-enforced single active lease per source and job type. Long commands refresh that same lease until they finish; they do not create disconnected start and finish records. Metadata is a structured JSON object, errors are a structured evidence array, and any skipped, quarantined or errored rows make the run partial rather than completed.

Automation is part of the evidence

06 · a scheduled job is only useful when it is distinct, observable and fail-visible

The production automation catalogue was reviewed end to end on 21 July 2026. Thirty-two distinct workflows remain across ingestion, verification, derivation, deployment and monitoring; five obsolete registrations are disabled. Jobs known to be blocked on hosted infrastructure are not presented as working cloud pipelines, and a red quality contract is kept red until the underlying data is repaired.

Least privilege
Read-only by default
Automation receives read-only repository access unless its job explicitly needs a narrower write capability, such as updating one deduplicated failure issue.
Fail-visible
Partial is not complete
Legacy scheduled importers are wrapped in the same required run lease as maintained importers. Health monitoring evaluates the latest result for the exact source, job and parameter contract, so a genuine rerun clears its own failure while an unrelated scraper_run cannot hide it.
Capacity-aware
Bounded concurrency
Parallel work is bounded to the capacity of the Data API, and long analytical scans use a direct database session so a gateway timeout cannot terminate valid server-side work.

Corrections in public

07 · trust is the moat

Every dataset has errors. Confirmed user-facing corrections are published in our public erratum log with what was wrong, what changed, when we found it and the verification source. We are also tightening this into a database-enforced correction ledger so operational repairs cannot bypass the public evidence trail.

A long erratum log is a good sign: it means we're looking. An empty one would be the thing to worry about.

Spotted something wrong?

Email hello@foot.io with the entity URL and the source you'd cite to correct it. We aim to publish a fix within a few business days — most ship same-day.

Read the erratum log →

What we don't do

08 · the honest defaults
No invented history
Empty > estimated
Advanced metrics like expected goals and spatial event data only exist for matches that were actually tracked. For older games those columns stay empty — we never back-fill a 1950s match with a modelled xG and present it as fact.
No silent edits
Changes are logged
User-facing corrections are published, and the engineering goal is a database-enforced audit trail for every factual mutation. Until that is complete, check the erratum and source metadata before citing a value.