How The News Directory works - sources, normalization and what we don't store

How this works

What we store, and what we don't

For each story we keep a headline, a short summary (capped at 200 characters - about the length of the blurb a publisher puts in their own page metadata), the outlet, the time it was published, and a link to the original.

We do not store or display article text. That is a deliberate property of the data model, not a habit: there is no field for an article body, so a source that hands us one has nowhere to put it. If you want to read a story, you click through to the publisher - which is the entire point of a directory.

One set of fields, whatever the source

Different providers describe an article differently. One calls it webTitle, another title; one gives a language and a country, another gives neither. We map each provider's own attributes into one normalized set of fields - source, section, language, country, topics, themes, keywords, publication time - and record which provider supplied each value, so a value is always traceable to where it came from.

Every filter and every sort option on this site is generated from that field set. That is why the same controls work on the front page, on a source page and on a theme page.

The same story, from several sources

When several outlets report one story we recognise it as one story - from the significant words in the headline, not from the URL, because aggregators rewrite those - and keep a single entry that records every provider which carried it. The Similar link on any headline shows you the same report from elsewhere, and the related coverage around it.

Themes and trending

Themes and keywords are our own parse of the headline and summary, not a provider's categories. Trending is a comparison, never a count: a term ranks by how much its share of the current window exceeds its share of the window before it, so terms that are always common do not dominate.

Dates

The date range sits in the top right of every board and defaults to the last 7 days, newest first. It is part of the URL, like every other filter, so any view you are looking at is a view you can bookmark or share.

Where do the headlines come from?

From open news sources whose own terms permit this use - currently the GDELT Project (licensed for commercial use with attribution), Hacker News, Wikimedia and the US Federal Register. Several well-known outlets are deliberately NOT included: their terms restrict their feeds to personal or non-commercial use, and we would need their permission first.

Do you show the full article?

No, and we could not: the data model has no field for article text. You get a headline, a short summary and a link to the publisher's own page.

How fresh is it?

Sources are polled about every 15 minutes. Where a source supports it we send a conditional request, so an unchanged feed costs one small response and nothing else - both faster and more polite.

Why does a source page sometimes look thin?

Because the feed behind it may have failed, and we say so on the page. A failing source is reported as a failure, never presented as a quiet news day.

Can I link to a filtered view?

Yes. Every filter, the sort order and the date range are all query parameters, and they update as you change them - so the address bar always describes exactly what you are looking at.

Something of mine is listed and I'd rather it wasn't.

We only ever hold a headline, a short summary and a link to your page - but tell us and we will remove it and stop indexing the source.