The primary source in the age of mechanical multiplication

“Many stories lay dormant in the vast amounts of data produced by everyday consumers because journalists are still only starting to acquire the large-scale data-wrangling expertise needed to tap them.”

Historians and journalists alike have long prized one source of information above all others: the on-the-record, primary source. They scour the attics for diaries and journals and fly across the country for interviews. But now we have a glut of documentation.

lam-thuy-voThanks to social media, millions of people go on the record, publicly, every single day. People sent billions of tweets, Facebook posts, and WhatsApp messages last year. They have expressed wonder 😮, anger 😤, and love ❤️ for social and political issues.

At present, most journalists treat social sources like they would any other — individual anecdotes and single points of contact. But to do so with a handful of tweets and Instagram posts is to ignore the potential of hundreds of millions of others.

Many stories lay dormant in the vast amounts of data produced by everyday consumers because journalists are still only starting to acquire the large-scale data-wrangling expertise needed to tap them. As more and more people conduct their lives online, and as smartphones are penetrating previously unconnected regions around the world, this trove of stories is only becoming larger.

The kinds of stories journalists can tell using this data are wide ranging. We can reconstruct online encounters in ways more precise than a source may recall from memory. Le Monde, for instance, retraced the journey of Syrian refugees through their WhatsApp messages. At Al Jazeera America, we analyzed and chronicled the evolution of a Hong Kong pro-democracy movement in Facebook chatrooms.

Journalists can try to find insight to people’s personality and character or hold powerful people accountable. Data scientist David Robinson did a sentiment analysis of Donald Trump’s tweets and found that Trump’s own tweets are much more negative than those his campaign staff tweeted. My colleague Charlie Warzel and I looked at the links Trump tweeted to explore the news he chooses to circulate, as a proxy for the news he may consume.

Journalists can examine the ways in which technology disadvantages groups by looking at social data. ProPublica’s Julia Angwin and Terry Parris Jr. bought a Facebook ad and established that the social media company allows advertisers to exclude customers based on race, while Vox’s Alvin Chang expanded on ProPublica’s analysis by looking at whether Facebook’s algorithm excludes already disadvantaged populations from being offered opportunities that their more affluent counterparts receive.

When journalists venture into this kind of story mining, I hope that they also continue to discuss the ethics surrounding it, the blurred lines between what is considered public and what is considered private, and the caveats that come with each dataset.

Last but not least, I hope that journalists will dig into social data to gain insights into whom they reach and, perhaps more importantly, whom they do not reach. People live in their filtered worlds in which algorithms serve them information that tends to affirm rather than question political views.

Maybe social data can allow us to understand these bubbles better in an effort to pierce them.

Lam Thuy Vo is a fellow in BuzzFeed’s Open Lab for Journalism, Technology, and the Arts.

Megan H. Chan   Cultural reporting goes mainstream

Nicholas Quah   Podcasting’s coming class war

Trushar Barot   API or die

Jon Slade   Trusted news, at a premium

Sarah Wolozin   Virtual reality on the open web

Laura E. Davis   Show your work

Aja Bogdanoff   Comments start pulling their weight

Erin Millar   The bottom falls out of Canadian media

Vivian Schiller   Tested like never before

Tim Herrera   The safe space of service journalism

Sydette Harry   Facing journalism’s history

Mark Armstrong   Time to pay up

Alberto Cairo   Communicating uncertainty to our readers

Gabriel Snyder   The aberration of 20th-century journalism

Eric Nuzum   Podcasting stratifies into hard layers

Samantha Barry   Messaging apps go mainstream

Carla Zanoni   Prioritizing emotional health

Libby Bawcombe   Kids board the podcast train

Margarita Noriega   From pinning tweets to tweeting pins

M. Scott Havens   Quality advertising to pair with quality content

Taylor Lorenz   “Selfie journalism” becomes a thing

Mary Walter-Brown   Getting comfortable asking for money

Pablo Boczkowski   Fake news and the future of journalism

Bill Keller   A healthy skepticism about data

Erin Pettigrew   A year of reflection in tech

Peter Sterne   A dangerous anti-press mix

Melody Kramer   Radically rethinking design

Ken Schwencke   Disaggregation and collection

Matt Karolian   AI improves publishing

Kawandeep Virdee   Moving deeper than the machine of clicks

Cindy Royal   Preparing the digital educator-scholar hybrid

Mathew Ingram   The Faustian Facebook dance continues

Maria Bustillos   “It’s true — I saw it on Facebook”

Mike Ragsdale   A smarter information diet

Cory Haik   Navigating power in Trump’s America

Christopher Meighan   Unlocking a deeper mobile experience

Kathleen Kingsbury   Print as a premium offering

Richard J. Tofel   The country doesn’t trust us — but they do believe us

Geetika Rudra   Journalism is community

Steve Henn   The next revolution is voice

David Weigel   A test for online speech

Alexis Lloyd   Public trust for private realities

Reyhan Harmanci   Bear witness — but then what?

Sam Ford   The year we talk about our awful metrics

Swati Sharma   Failing diversity is failing journalism

Amy O'Leary   Not just covering communities, reaching them

Felix Salmon   Headlines matter

Jim Friedlich   A banner year for venture philanthropy

Katie Zhu   The year of minority media

Burt Herman   Local news gets interesting

Rubina Madan Fillion   Snapchat grows up

Francesco Marconi   The year of augmented writing

Umbreen Bhatti   A sense of journalists’ humanity

Laura Walker   Authentic voices, not fake news

An Xiao Mina   2017 is for the attention innovators

Robert Hernandez   History will exclude you, again

Almar Latour   Thanks, #fakenews

Ashley C. Woods   Local journalism will fight a new fight

Dannagal G. Young   The return of the gatekeepers

Juliette De Maeyer and Dominique Trudel   A rebirth of populist journalism

Liz McMillen   The year of deep insights

Valérie Bélair-Gagnon   Truthiness in private spaces

Millie Tran   International expansion without colonial overtones

Olivia Ma   The year collaboration beats competition

Keren Goldshlager   Defining a focus, and then saying no

Tanya Cordrey   The resurgence of reach

Asma Khalid   The year of the newsy podcast

Emily Goligoski   Incorporating audience feedback at scale

Sue Schardt   Objectivity, fairness, balance, and love

Adam Thomas   The coming collaboration across Europe

Errin Haines Whack   Chaos or community?

Rachel Schallom   Stop flying over the flyover states

Jeremy Barr   A terrible year for Tiers B through D

Liz Danzico   The triumph of the small

Jonathan Stray   A boom in responsible conservative media

Ernst-Jan Pfauth   Earn trust by working for (and with) readers

Emi Kolawole   From empathy to community

Renée Kaplan   Pure reach has reached its limit

Ståle Grut   The battle for high-quality VR

Molly de Aguiar   Philanthropists galvanize around news

Dan Colarusso   Let’s make live video we can love

Dan Gillmor   Fix the demand side of news too

Javaun Moradi   What can we own?

Michael Kuntz   Trust is the new click

Michael Oreskes   Reversing the erosion of democracy

Nushin Rashidian   A rise in high-price, high-value subscriptions

Guy Raz   Inspiration and hope will matter more than ever

Nathalie Malinarich   Making it easy

Julia Beizer   Building a coherent core identity

Matt Waite   The people running the media are the problem

S.P. Sullivan   Baking transparency into our routines

Coleen O'Lear   Back to basics

Jonathan Hunt   Measurement companies get with the times

Joanne Lipman   The year of the drone, really

Zizi Papacharissi   Distracted journalism looks in the mirror

Tracie Powell   Building reader relationships

Ryan McCarthy   Platforms grow up or grow more toxic

Andrew Haeg   The year of listening

Alice Antheaume   A new test for French media

Carrie Brown-Smith   We won’t do enough

David Skok   What lies beyond paywalls

Mary Meehan   Feeling blue in a red state

Amie Ferris-Rotman   Вслед за Россией

Priya Ganapati   Mobile websites are ready for reinvention

Bill Adair   The year of the fact-checking bot

David Chavern   Fake news gets solved

Elizabeth Jensen   Trust depends on the details

Andrew Losowsky   Building our own communities

Ray Soto   VR moves from experiments to immersion

Andrew Ramsammy   Rise of the rebel journalist

Mira Lowe   News literacy, bias, and “Hamilton”

Doris Truong   Connecting with diverse perspectives

Andrea Silenzi   Podcasts dive into breaking news analysis

Juan Luis Sánchez   Your predictions are our present

Tim Griggs   The year we stop taking sides

Andy Rossback   The year of the user

Annemarie Dooling   UGC as a path out of the bubble

Rasmus Kleis Nielsen   News after advertising may look like news before advertising

Rebekah Monson   Journalism is community-as-a-service

Ole Reißmann   Un-faking the news

Scott Dodd   Nonprofits team up for impact

Helen Havlak   Chasing mobile search results

Sara M. Watson   There is no neutral interface

Mandy Velez   The audience is the source and the story

Mario García   Virtual reality on mobile leaps forward

Lam Thuy Vo   The primary source in the age of mechanical multiplication

Caitlin Thompson   High touch, high value

Dhiya Kuriakose   The year of digital detoxing

Amy Webb   Journalism as a service

Corey Ford   The year of the rebelpreneur

Ariane Bernard   Better data about your users

Claire Wardle   Verification takes center stage

P. Kim Bui   The year journalism teaches again

Sarah Marshall   Focusing on the why of the click

Tressie McMillan Cottom   A path through the media’s coming legitimacy crisis

Lee Glendinning   A call for great editing

Anita Zielina   The sales funnel reaches (and changes) the newsroom

Hillary Frey   Forests need to burn to regrow

Moreno Cruz Osório   The year of transparency in Brazilian journalism

Rachel Sklar   Women are going to get loud