AI Agent
Entity Mapping Agent
Turns a raw news article into correctly tagged, deduplicated entities the rest of the platform can query, the invisible work behind every "who's this about" answer.
What is it?
A two-flow pipeline that runs on every ingested article: first extracts every named entity with its type, sport and confidence, then, for anything that doesn't match an existing record, makes a second-pass judgment call on whether it's a known alias or a genuinely new entity.
What does it do?
- -Two models, two jobs: a fast, cheap model for high-volume extraction, a second model reserved for the harder disambiguation decisions.
- -Never blocks ingestion: if entity resolution fails on one article, the article still ingests, it just stays unlinked rather than failing the whole pipeline.
- -Grows its own reference data: every confirmed new entity or alias gets written back, so future articles resolve faster.
- -Real scale today: 102,035 canonical entities and 10,954 known aliases, the entire set held in memory and refreshed every 15 minutes.
Why does it matter?
An LLM extracting names from text will find candidates, but deciding whether "Man Utd" and "Manchester United" are the same entity is a judgment call, not a lookup, and it has to happen at ingestion time or every downstream feature inherits the mess.
Who is it for?
Engineering
How it works
Flow 1: a fast, cheap model extracts named entities from article text via the AI Gateway. Flow 2: for unmatched surface forms, Claude Haiku judges each against up to 20 candidate matches and decides alias versus new entity, batched to stay well inside context limits.