Nieman Foundation at Harvard
HOME
          
LATEST STORY
The New York Times’ new Slack 2016 election bot sends readers’ questions straight to the newsroom
ABOUT                    SUBSCRIBE
June 17, 2009, 2:01 p.m.

Knight News Challenge: A grant to DocumentCloud promises a data boost for investigative journalism

The Knight News Challenge‘s biggest winner, with a two-year grant of $719,500, is DocumentCloud, the primary-source index conceived by journalists and developers at ProPublica and The New York Times. Here’s why you should care: There’s good reason to believe the project will transform how some investigative journalism is conducted — and who conducts it.

Like a lot of software in the cloud, this one is complicated to explain. I wrote a long overview of DocumentCloud in November, and you can read their initial grant application in my first post about the project. Aron Pilhofer, editor of interactive news technologies at the Times and one of the project’s creators, told me on Monday, “DocumentCloud isn’t really conducive to a two-minute elevator pitch.” But later in our conversation, he ventured one: “It will turn documents into data.”

In the analog version of investigative journalism, a reporter obtains documents from sources and freedom-of-information requests, writes a story, and… that’s it. If we’re lucky, the materials are posted as unwieldy and barely searchable PDFs.

DocumentCloud’s vision is to collect, archive, and index the text and metadata of all documents used by participating news organizations, advocacy groups, bloggers, and others — “so they’re not just sitting in the corner of a newsroom collecting dust,” Pilhofer explained. That way, anyone — from other news outlets to curious readers — will be able to search across all documents in the project to find information that might not have been relevant to the original piece. If it were an animated TV series, the catchphrase might be, With our newsrooms combined — we are DocumentCloud!

Early partners in the project include the Times, ProPublica (the non-profit investigative journalism outfit) Gotham Gazette (a New York City news site published by Citizens Union Foundation, themselves winners of two Knight News Challenge grants), TPM Muckraker (the investigative arm of Talking Points Memo), and the National Security Archive (home to the largest public repository of declassified government documents). Are you salivating yet?

Anyone who has waded through the National Security Archive’s wealth of FBI files and CIA reports will immediately recognize the benefit of DocumentCloud. What if you could search across the entire archive for a particular topic of interest (Pilhofer suggested Marilyn Monroe) and get pinged whenever that topic shows up in a new document? Or what if you were a business journalist up to your neck in SEC filings? Or a local blogger keeping tabs on your congressman’s earmarks?

Details are still being worked out, but running materials through DocumentCloud will involve some sort of optical character recognition (to render those pesky image-based PDFs favored by some government agencies into searchable text) and Open Calais (to extract metadata like names, locations, and dates for more effective indexing). The news organizations that contribute documents will likely host them as well, said Scott Klein, editor of online development at ProPublica and a creator of DocumentCloud. He described the project as more of a “card catalog” than, say, a repository.

There are other aspects of the project that I’d be happy to discuss in the comments, and maybe we can get Pilhofer and Klein to weigh in here as well (as though I’m not hyping them enough). The other creators are Eric Umansky of ProPublica and Ben Koski of the Times. So if you have any questions about DocumentCloud, feel free to ask, and I’ll work on getting answers.

UPDATE, 3:30 p.m.: The Document Cloud folks introduced their project at the Future of News and Civic Media Conference at MIT today, and I shot some raw video with Qik:

POSTED     June 17, 2009, 2:01 p.m.
PART OF A SERIES     Knight News Challenge 2009
SHARE THIS STORY
   
Show comments  
Show tags
 
Join the 15,000 who get the freshest future-of-journalism news in our daily email.
The New York Times’ new Slack 2016 election bot sends readers’ questions straight to the newsroom
“Instead of asking you to come to us and be part of this massive room of people shouting over each other, you can bring us to you, and have us be, essentially, one more person in your conversation.”
The Conversation expands across the U.S., freshly funded by universities and foundations
The news site that uses academics as reporters and journalists as editors now boasts 19 paying member universities and is opening up posts in Atlanta (and maybe in the Bay Area).
A Boston public radio station is redesigning its site to make audio “a first-class citizen online”
But: “I’ve tried to be really disciplined about not calling this process just a redesign,” WBUR’s executive editor for digital Tiffany Campbell said. “We’ve built a brand new platform.”
What to read next
0
tweets
Newsonomics: Setting the news table for 2016
The news business hopes it won’t end up one sandwich short of a picnic as the new year’s big trends unfold.
0The sun never sets on The Times: How and why the British paper built its new weekly international app
“We’re pursuing the idea of editions everywhere. An edition is something that can be finished. When you’ve read it, you feel up-to-date; you’ve been told what you need to know for the day or the week.”
0Hot Pod: Is the next front in podcast innovation hardware?
Plus, 21st Century Fox invests in a new podcast network, and some thoughts on the second season of Serial.
These stories are our most popular on Twitter over the past 30 days.
See all our most recent pieces ➚
Encyclo is our encyclopedia of the future of news, chronicling the key players in journalism’s evolution.
Here are a few of the entries you’ll find in Encyclo.   Get the full Encyclo ➚
The Daily Show
Outside.in
Bureau of Investigative Journalism
The New Republic
The Nation
Las Vegas Sun
CBS News
Conde Nast
The Fiscal Times
Newsweek
Placeblogger
Texas Tribune