Top
Navigation
Excellence in AI Innovation, Excellence in AI Innovation, Large Newsroom finalist

The New York Times’s A.I. Toolkit for Investigative Journalism


About the Project

We’ve developed a suite of A.I.-powered tools and processes that allow reporters to interrogate vast, messy datasets while keeping primary documents and human judgment at the center of the work.

The backbone of this entry is Cheatsheet, an internal web app that looks like a spreadsheet but is powered by “recipes” — vetted A.I. workflows for investigative tasks such as quote extraction, summarization, translation, web search and information classification. Reporters upload datasets ranging from public records and FOIA document dumps to video transcripts and social media posts, then chain recipes together to score, group and transform the material into structured, reviewable rows and columns. Cheatsheet has now supported dozens of investigations across the Times, and is empowering hundreds of journalists of all technical levels to run sophisticated analyses.

In The Epstein Files, the newsroom confronted roughly 3 million pages, 1.1 million emails and 200,000 images and videos released from the Justice Department’s disclosures about Jeffrey Epstein and his network. The Times augmented DocTools — its in‑house document platform — with new A.I.-driven tools that allowed reporters to semantically search images and videos. We also built the “Epstein Files Engine,” an A.I. agent that responded to any reporter question about the corpus. Powered by our duplicate-detection algorithm, it contextualized findings with what The Times already knew and reported to deliver citation-rich leads about key individuals and storylines in mere minutes.

The same A.I. investigative toolkit underpins several marquee investigations included in this entry. To probe Donald Trump’s age and cognitive fitness, reporters used Cheatsheet to score more than six million words of speeches for themes like rambling, fixation and profanity, then distilled those results into a tightly vetted set of examples that undergirded a narrative about how his rhetoric has shifted over a decade. To investigate Musk’s Grok chatbot flooding X with sexualized deepfakes, Times journalists collected 525,000 images and ran two computer vision models — one to detect women, another to detect sexualized content — to demonstrate that at least 41 percent of posts likely contained sexualized imagery of women.

For the Satoshi Nakamoto investigation, A.I. helped examine hundreds of thousands of mailing-list posts, surfacing distinctive grammatical tics and “synonym‑less” technical terms that pointed to a single suspect, analysis that was then rigorously double‑checked against the underlying corpus. And in “Chatbots Can Go Into a Delusional Spiral,” reporters used A.I. to map themes across more than 3,000 pages of transcripts with ChatGPT and to demonstrate that other chatbots exhibited similar sycophantic tendencies when presented with the same scenarios.

Across all of this work, A.I. is treated as an assistive layer for search, extraction and scoring — never as an author. Reporters are trained to treat outputs as leads, not facts, and journalists check every quote or finding against the primary evidence before it’s published. The result is a newsroom-wide A.I. practice that enables ambitious, data‑driven investigations that would be impractical or impossible otherwise, while keeping the reporting grounded in human expertise.