Storytelling with graphics: From raw data to reader impact

By • Observable webinar in Zoom

Recording

Show the timestamped transcript
  1. All right. Hello, everybody, and welcome. So we at Observable are so excited to host
  2. this discussion on how data journalists approach storytelling with data. So in a moment I will
  3. be introducing our panelists, but before I do, I just want to go through some quick logistics.
  4. So first of all, one of our panelists, Jackie Schrag, unfortunately will not be able to
  5. participate today, but we do hope she can join us in a future session. Now, what to
  6. expect today. So in just a moment I'll introduce our panelists and we will dive into our conversation,
  7. which we expect to run for about 30 minutes. After we wrap up the roundtable, we're going
  8. to have about 10 minutes of Q&A, and you're welcome to either add your questions in the
  9. chat during the roundtable at any point, or you can wait until the Q&A to do so. And we'll try to
  10. get through as many questions as we can in that time. Also, this webinar will be recorded, so
  11. don't worry if you have to hop early or if you want to revisit it in the future. All the
  12. registrants will get access to the recording, so you can always come back to it. Now, without further
  13. ado, let's welcome our panelists to the stage. So welcome, Ben, Kavya, and Jared. First, let me
  14. do a quick introduction of myself. So I am Will Chase. I am a senior product designer here at
  15. Observable. But prior to joining Observable, I did work in data journalism for about four years.
  16. I used to run the visual storytelling team at Axios, so returning to my roots here.
  17. So first, I want to welcome Ben Welsh to the stage.
  18. Hello. Welcome, Ben. Hello, hello. So Ben is a reporter, editor, and computer programmer. He
  19. is the founder of the Reuters News Application Desk, which covers the world's most pressing
  20. stories by developing dashboards, databases, and other automated systems. Welcome, Ben.
  21. Thanks for having me. Great. Next, I'm going to welcome Jared Whalen.
  22. Hello, Jared. Okay. So Jared is a Philadelphia-based journalist engineer
  23. at Polygraph and The Pudding. Previously, he's worked at Axios, the Philadelphia Inquirer,
  24. and the Delaware News Journal. His background includes visual storytelling,
  25. web application development, and data journalism.
  26. All right. And finally, let me welcome Kavya Beharaj to the stage.
  27. Hey. Hello. Welcome, Kavya. So Kavya is a designer and developer based in Brooklyn.
  28. She is currently an associate editor of data visualization at Axios. And before that,
  29. she worked in the nonprofit sector managing data projects, building web apps, and training others
  30. to better understand and use their data. All right. So thank you so much, everybody,
  31. for being here. We can go ahead and jump right into our discussion questions. So first, we are
  32. going to do a little bit of show and tell. So I'm going to ask each of our participants to share
  33. what they felt was one of the most impactful visualizations that they have made or worked on.
  34. And talk a little bit about it. Talk about, you know, what they think made it so impactful.
  35. So let's go ahead and start with Ben. Go for it, Ben.
  36. Hey. So I'm going to share my screen, and we're going to step into a time machine
  37. back to the year 2012. Who remembers that? And this is a graphic I worked on at the time with
  38. Robert Lopez and Kate Lithicum, two of my colleagues then at the Los Angeles Times
  39. in Southern California. And we put together this map, which is called How Fast is LAFD,
  40. where you live, to chart the response times to 911 calls by the local fire department
  41. there in the city. And you see it visualized there on the map, that sort of sideways key
  42. shape, believe it or not, is the actual city limits of Los Angeles itself. And it's surrounded
  43. by 87 other suburbs in LA County. And so if you ever see a weird shape like that, that's
  44. just because that's how the city was drawn up more than 100 years ago. And that is the
  45. jurisdiction of the city fire department that respond not just, of course, to structure fires
  46. and wildfires within the city limits, but primarily to emergency medical care situations.
  47. And there had emerged at the time, a sort of political controversy on a then very popular
  48. local news website called the Huffington Post, where a candidate for mayor who is trying to
  49. run for mayor again now in 2026, believe it or not, made some allegations about the fire department
  50. being substandard and responding to 911 calls. And that created a little bit of a political
  51. controversy and kind of raised the question within the politics of the city about what's
  52. going on at the fire department. And then curiously, the fire department invited me and
  53. Kate Lint to come over one day because we were asking, what's up? You got any data? Let's try to
  54. answer the questions from this Politico. And they called us over to City Hall East, if you've been
  55. there in downtown LA and put us in a conference room and they said, well, we got to tell you guys
  56. something. All the statistics we've been publishing for the last five or two years or so,
  57. they're all wrong. You know, the guy who had been our stats guy, well, he retired and it just kind
  58. of been on autopilot and we were looking at him and we realized he just did the math wrong.
  59. It was one of the stranger sort of moments as a reporter to just like be invited in for a
  60. confession, you know, and I think it really was kind of a well-meaning moment. They were trying
  61. to be transparent, but it's also pretty embarrassing, right? And so the front page
  62. of the newspaper the next day had a story by Kate and I that basically said, look, they say
  63. they don't know, right? And that just made the question an even bigger question. What the heck
  64. is going on with 911 at the city? This candidate had made very specific allegations. He couldn't
  65. really back him up and we couldn't, the department didn't really have an answer. So we filed a public
  66. records request for the complete database of millions and millions of 911 calls to try to
  67. answer the question ourselves, which we felt was kind of now in the public interest. And there's
  68. a whole long story I won't bore you with about fighting for the FOIA and getting the data and all
  69. the characters and blah, blah, blah. But at the end of the day, we kind of did our own analysis of
  70. those millions of calls and we used an outside standard, which is always a really powerful way
  71. to structure any kind of accountability work you do with data. So you see color coding on this map
  72. where green is good and purple is bad, pretty standard. But those thresholds and those cutoffs
  73. and the colors were all based on sort of the national response time standards of how fast
  74. calls are supposed to be. And that was a standard that the fire department itself had embraced and
  75. said it was doing well on. And so we were able to bring to bear sort of an outside standard that
  76. allowed us to evaluate, you know, kind of transparently and clearly and with, you know,
  77. colors on a map accountability or bring accountability to the question. And, you know,
  78. in any data visualization, if you want to get more investigative, you know, having that outside
  79. standard or something to measure up against that isn't just like according to you or according to
  80. its standard deviation in the distribution, but some real benchmark, right? So we had that strong
  81. benchmark and then it turned out, believe it or not, we had a reveal, you know, and that's a bit
  82. of Reuters ease. That's the term we use here at Reuters. When your report is going to tell the
  83. world something, it doesn't already know. And to me, that's the second piece that makes any powerful
  84. investigation really break through. And, you know, the reveal really was that the politicians claims
  85. were all wrong. You know what I mean? But that really was when you looked at the data closely,
  86. there were two or three really significant weaknesses and how 911 care is being delivered
  87. that were not being discussed in the city like at all. Like we're not even on the agenda. No one
  88. even talked about them. And they really surfaced and came up out of the data. And then we use that
  89. data to follow it out into the world and report on real world cases and show the impact and stakes.
  90. And so this visualization ran online as an interactive map, allowing people to personalize
  91. the story, which, of course, is another classic way to bring data out in journalism to look up
  92. their neighborhoods, but also to kind of tell the overall story, which ran on the front page of the
  93. print edition the same day. There's the print map that my colleague Tom Lauder put together
  94. with slightly different styles, but kind of the same idea. And you can see that the first
  95. angle we ran with the first day, and you might have to know LA to pull this out of the map,
  96. was that it was actually the richest neighborhoods in the city had the slowest response times,
  97. which is the Hollywood Hills where all the fancy movie stars live up at their mansions next to the
  98. Hollywood side. And the wealthier areas had the slowest response because they're in rural hillside
  99. communities is like the fundamental reason. But that was kind of the opposite of the way people
  100. had been talking about the issue. And so that was kind of a way into it that was kind of a reveal
  101. that gave you, hey, the truth of it is that it's actually the rich people who get the slowest
  102. response on average, and we built the whole story around that and told some real life cases of where
  103. people had died when calls had gone wrong in those areas. But if you look at the map closely,
  104. you'll see some other interesting things too, which is like, and this really stuck out to me
  105. at the time, is like, look, it's like the edge of the city on the borders is where it's really slow.
  106. What's up with that? No one talks about that. Well, we looked into it and it's common sense
  107. when you think about it, the fire trucks all sit at the fire stations. And if you're going to build
  108. 108 fire stations, you don't put them next to the city border because you want to put them where
  109. they have a radius around them that they're going to serve. But then geospatially, you guys are all
  110. smart and dated people here, you have a weakness. You're then slower to get anywhere near the
  111. borders and particularly that long thin strip of the city that runs down to the harbor. And so we
  112. did a second story that was about how these boundaries and the lack of connections with
  113. neighboring fire departments to pool services was another big issue with the system. And then we did,
  114. this is a boring story I wrote because we needed to have one that day. But then we did a fourth
  115. one that was about how actually when you looked at the segments of the call, from when they pick
  116. up the phone to when someone arrives at your door, there's two or three segments in there of like
  117. triaging what happened, sending somebody getting to the place. We found that that early call taking
  118. period was actually where the biggest lag was in the whole system. And we were able to focus on
  119. that discrete flaw. And we did a whole lot of other coverage and lots and lots happened. And
  120. it was all tied up in city politics and an election that happened. But the result is,
  121. is the next mayor made really deep and significant changes to how the 9-1-1 system works,
  122. including revising how the call taking is done in the calling scripts. And also established a
  123. new statistical unit in the fire department to do the type of analysis we were doing internally
  124. within the department for itself. They called it fire stat. And the fire chief was replaced by,
  125. as well as part of the new administration. And then that ultimately kind of sent me down this
  126. road of becoming a fire department reporter, though never really planning to, which yielded
  127. later investigations about a failed building inspection program that led to overhauls there.
  128. And we were able to uncover and expose nepotism and abuse of the hiring process that led to
  129. complete reform of that as well and a few other things. And it all started with trying to make
  130. this map on my laptop in like summer of 2012. And so that's my long and boring tale. I hope it
  131. wasn't too long. No, not at all. I was going to say not boring in any way. I think it's actually,
  132. that's a really great story. Every step of the process, I was going to say,
  133. oh, and so which areas are the green areas that you noticed? Is it just because it's
  134. harder to get to? And then what was the result? Did they overhaul the things? So yeah,
  135. you answered all the questions. Yeah, the green areas are where the fire stations are.
  136. And so many other fire services will preposition units in high volume call
  137. zones rather than have them sit at the station. That's like one common response to that issue.
  138. The fire department has not done that in Los Angeles. And if you paid close attention to
  139. the Palisades fire earlier this year, one of the key criticisms of how that was managed is they
  140. did not preposition units in areas that were prone to fire, even though there was a very high risk
  141. fire warning at the time. My colleagues who are still at the LA Times did some great reporting
  142. showing how that structural flaw still exists. Yeah, because it does really look like it's
  143. essentially gridded across the city like that. Yeah. The other thing I think is amazing,
  144. I'll just call out is that you mentioned this was made in 2012, and it is still running on your
  145. browser here more than more than 10 years later. So it's always nice to see the longevity of things
  146. like that. I'll publish static files if you can, guys. That's my static files. There you go.
  147. All right. Jared, why don't you go? Why don't you go next?
  148. Okay, certainly. Hi, again. Hi, everybody. My name is Jared from Philadelphia. The project I
  149. want to show today is one that over at Polygraph, which is the internal studio of The Pudding,
  150. which I hope many of you are familiar with, a project that we recently published that I thought
  151. kind of covered a lot of different elements of being a data journalist. For this project,
  152. we partnered up with the Trans-Journalist Alliance Association to take a look at the question of
  153. how can we kind of audit the news industry and look at how they are exploring a particular topic,
  154. in this case, trans issues. The way we went about doing that is we put together a script that is
  155. using the Media Cloud API, which is an organization that basically compiles news articles by
  156. keyword search and just allows you to, while not actually pull the body text for obvious reasons,
  157. allow you to say, I want to see all the articles that include these keywords. I want to pull all
  158. the articles from the New York Times that mention this topic. From there, we just put together a
  159. long list of keywords, a long list of things that we felt captured the breadth of trans issues
  160. across healthcare, across pop culture, across social issues, and then we're able to generate
  161. this corpus of hundreds of thousands of articles over the course of the last five years. From there,
  162. we were then able to run the data through a large language model classifier to then take those
  163. articles and group them by themes that we predetermined. For example, we were able to
  164. take all those articles, filter out the noise, and then come out with 18,000 articles about
  165. healthcare and body autonomy. From there, we then were able to then go into each of these
  166. broader themes and cluster those articles by events. From here, we're not only looking at
  167. the 18,000 articles, but we were able to find those 150 events that group those articles together.
  168. And why do all this? Well, the reason we wanted to do this is because we wanted a way
  169. to hold ourselves in the news industry accountable for the stuff that we're covering. How are we
  170. talking about these issues? So in this case, when you look for this broader theme, we can see some
  171. of the major stories that pop out, such as convergent therapy going before the Supreme Court
  172. or puberty blocker bans. And what we hope to do is create an interactive experience that allows
  173. the user to really dive in, look at these individual articles, explore how different
  174. publications based off maybe how their political lean are, and see how they are covering it.
  175. And then what I really enjoy is because we partnered with this organization, we were able
  176. to bring in their insights. So we were able to have them come in and say like, well, here is how
  177. a topic is being covered. And here are some insights, how we think maybe this topic should
  178. be covered, or how us in the trans community believe that it should be covered. From there,
  179. we basically present the data in ways that we hopefully communicate the story, such as
  180. publication, political lean by year, giving the articles themselves, showing these little
  181. bee swarms for each topic of how, based on your current filtering, how the articles are published,
  182. which allows you to see jumps across time when a story is in the news. And what we want to hopefully
  183. do with all this is as this project continues to grow, because we have this running on an ongoing
  184. basis, is be able to take a look back at any given point and look for those stories, look for
  185. how, as the political climate changes and things in both the news, but as a way of us in the news
  186. industry, change how we think about toppings and how to cover things, be able to actually put some
  187. data behind some of those preconceived notions. The response to this project has been
  188. overwhelmingly successful, which we always want to hear. And my hope is that as we continue to
  189. maintain this project over the coming years, that we're able to see the news industry in general do
  190. a bit more inward looking, but taking, investing the time to do audits of your own reporting,
  191. your own coverage to see how you are reporting on the communities you read you.
  192. Wow, that's amazing. Thanks, Jared. So when you say it's updated, do you mean that this is a
  193. sort of semi-manual process that you guys are returning to, or it's more of an automated
  194. updating process? Yeah, great question. So the way it works now, we have a monthly script that will
  195. one, pull the entire corpus of articles and then run it through our model to handle the classification.
  196. But as anybody who's worked with large language models, who's worked with machine learning,
  197. you know that when you're working with machine learning, you need to constantly training your
  198. model. So the way that we have this is we have a handful of configuration files that as new data
  199. comes in, it gives us the ability to manually audit, to manually make corrections, and then
  200. continue to train the model going forward. So big picture of the way this works is on a monthly
  201. basis, it's doing classifications based off the existing data. And then we have a more semi-annual
  202. approach to kind of look at the entire corpus of data for reclassification purposes.
  203. Awesome. All right, let's go ahead and move on to Kavya.
  204. Kavya, I think you're muted actually. My bad. Oh, good. Okay, so this is a
  205. project that I worked on with Will and Jared actually, and a couple other folks on our team.
  206. And this was right after that summer with intense wildfire smoke on the east coast.
  207. So I in New York was looking outside my window and seeing just orange. So what we wanted to
  208. accomplish with this article is basically kind of recap what happened that summer, what made
  209. the wildfire smoke that summer so unusual, and what kind of it means if this is our new normal
  210. with pollution exposure. So this was a collaboration with a lot of different people.
  211. You know, but the part that I want to zoom in on is the zoom in section that focuses on
  212. kind of like Ben was saying, localizing the story to you. So this is basically a dashboard that
  213. lets you search for some metro areas and look up how your pollution exposure might have been
  214. in that time period compared to the average over the last, you know, 10 or so years.
  215. And this is a part of the motivation for doing this is because Axios has many local bureaus.
  216. So we have newsletters that operate in more than 30 US cities. And so we wanted a way to
  217. take this really large national story and make it relevant in some way to folks. So this dashboard
  218. and the data that I collected as part of it was for that effort. And also, you know,
  219. including some tips for how to actually protect yourself from wildfire smoke because that's
  220. something we need to do now. So part of the design decisions that went into this
  221. was kind of trying to make a chart that can kind of speak on its own terms.
  222. So if you just look at it, it tells a complete story. This is the most recent data. This is
  223. where this is how it compares to the average, and a pretty clear indication of whether what
  224. you're experiencing is healthy or unhealthy, according to the standard set by the government.
  225. So one of the effects or the impacts of like taking the time to scrape this data and clean
  226. it and prepare it is that our local markets were able to run stories based on it. So it wasn't just
  227. published once. The stories got used a lot after that too, because we took the time to localize it.
  228. So yeah, that's all. Also, the folks who built this are also on the call. So if you want to
  229. say anything about it as well, you both are welcome. I can do a quick answer to this question
  230. that Amanda just asked in the chat. How did you create the smoke animations on scroll?
  231. So yeah, that was another collaboration. A lot of people worked on this section.
  232. Essentially, this is a background image, which is the map. So it's a static image. So that map was
  233. created in a combination of tools, QGIS and Adobe Illustrator. And then we pulled the data for
  234. how much smoke there is. So NOAA publishes satellite data that you can process to
  235. get this density of smoke. And then we processed those images to create transparent versions that
  236. just have the brown portion where it's the intensity of smoke and the rest of the image
  237. is transparent. And so then those are overlaid on top of that background map image.
  238. And Jared worked on a
  239. scrolly telling a little bit of code that would cycle those images. So it steps through all the
  240. smoke images. We have like one image, I think, for every four hours, I want to say. I don't
  241. remember if that's exactly right. But it steps through, you know, all the hundreds of images
  242. for the smoke that we have that represent that time period. Yeah.
  243. All right. Thank you, everybody. So moving on now into a couple more questions just about how
  244. things work in newsrooms, how you deal with data. And anybody, feel free to jump in here.
  245. So when you're handling a messy data set, what is your first step in figuring out
  246. if there's a story worth telling? I can start. I think that a good first step, or at least
  247. what should be a good first step, is when you get a messy data set, or really any data set,
  248. but especially one that's messy and a bit complicated, is before you jump in to try
  249. to tell the story with it, before you jump in to try to find a narrative, is you make sure you
  250. figure out why this data set even exists and how it's already being used. It's very tempting to
  251. dive right into it and inject your interpretation of what it could tell you, especially if it is
  252. messy and maybe it presents some like story agnostic challenges like pulling from a PDF or doing some
  253. weird scraping, because a lot of us in this field see that and while it may be tedious, we actually
  254. kind of get excited about figuring out that problem. But what can often end up happening
  255. is you spend a lot of time to get the data to where you think it should look like, only to find
  256. it doesn't actually tell you what you think. So whether it's referring to documentation, whether
  257. it's seeing how other news outlets or other places are using that data, or hopping on the phone and
  258. talking with the data manager, those are pretty much what should be your first step anytime you
  259. encounter a new data set. Yeah, and I would say another technique is to just really focus on the
  260. fundamentals of the table itself. What is a row? And I mean that in the most philosophical
  261. and Socratic sense, you know what I mean? Like said, what does each row represent? What is it
  262. recording, right? What does each column indicate? Like I had an interesting example with some
  263. students I was working with a couple years ago at DePaul University, where they had been given a data
  264. set from the city of Chicago that was trees planted by the city tree planting program. We were going
  265. to look, you know, they had set a target for how many they were going to grow. There's your benchmark
  266. guys. And we were going to map it to see who got the trees. It's almost like I do this same story
  267. over and over again. And we did our provisional analysis, and went back to the city to kind of
  268. tell them what we had found. And we learned something a little embarrassing, which was that
  269. there was a column with a very cryptic name that we didn't know what it was, and didn't think too
  270. much about. But it was actually the number of trees that had been planted in that row. So each
  271. row was not a tree, which is what we were working under the assumption of. Each row was a work order
  272. to plant one too many trees, right? And so our initial analysis was kind of fundamentally wrong,
  273. because we didn't understand, you know, everything we needed to about the data. What is a row,
  274. and what does each column mean? And just slowing down and forcing yourself to walk through that
  275. is, I think, crucial. Yeah, I think my thoughts on that were pretty much covered. But I think
  276. there is also a way that focusing on one row of data can also make clear what steps are needed
  277. to clean it. You know, as you're trying to understand what is actually being shown, maybe
  278. it's multiple things at once. It's not just one observation, you know, maybe there's multiple
  279. data sets mixed together, they all need to be teased apart. And that's another thing that can
  280. come out of just trying to understand the data on a basic level. And you can trust your gut,
  281. if something is confusing to you, then it's probably going to be confusing to other people.
  282. Yeah, that's great. Jared, I love your point about like, just pick up the phone and call somebody,
  283. because, you know, I think so many of us in the data world or who come from a background that
  284. might be more like engineering focused. That's just like never where your mind would go. At first,
  285. right? Your mind goes to like, Oh, what code can I write to do this? Or how can I investigate this?
  286. And I remember there's a hilarious quote, Hadley Wickham, who, you know, works at our studio,
  287. he went to a journalism conference one time. And he said that after talking to people, he said,
  288. the thing I learned the most is that like, what you can learn by just picking up the phone
  289. is incredible. And that is a thing that I would never, literally never do in my life.
  290. So yes, pick up the phone, don't be afraid of it.
  291. Kavya, I actually wanted to come back to you for the next question. I was,
  292. because I know that you've had some experience working in other industries, you mentioned you've
  293. worked in nonprofits, and done some, you know, sort of dashboarding work in in the business space.
  294. So I know a lot of people in our audience create visualizations for businesses or internal teams,
  295. dashboards, that sort of thing. And what principles do from data journalism, do you
  296. think that they can apply to make their own work more impactful? Of course, anyone else is welcome
  297. to jump into it. But yeah, yeah, totally. I think in my past job, I've had a lot of experience
  298. making dashboards that I'm not sure anyone actually used or saw. And I think part of the
  299. challenge there is that we're going in assuming that we need a dashboard, you know, and that all
  300. of the data points that we want to visualize matter. But I think there are more fundamental
  301. questions, you know, like, why do we need to visualize all this in this way, who is going to
  302. be using it and how, and how is it going to be maintained, which is was my job before this.
  303. And so I think that those are principles from data journalism that apply anywhere.
  304. Like anytime you want to do some sort of visual communication or chart, it is really helpful to
  305. understand who you are trying to help, to influence, to inform. And making something
  306. that's actually useful starts with figuring out why it exists. That's my take.
  307. Yeah, that's great. I love that. All right, let me let me ask for the room a little bit about
  308. about collaboration. So a lot of the pieces we showed here in the in the sort of show and tell
  309. section, you know, there were many names on those or lots of people that were involved in those
  310. projects, you know, you've got reporters, designers, editors, developers, all kinds of people, artists
  311. working on them. So what have any of you learned about working across roles that might help other
  312. people building data visualizations in collaborative environments?
  313. I think something that I learned, especially working at Axios is just the importance of
  314. like that it's okay to let go. I know a lot of newsroom developers, but I imagine a lot of the
  315. people who work in smaller companies as well. You're probably one of the few technical people
  316. on your team, you might very well be the only designer, the only developer, the only data
  317. journalist in your newsroom in your organization. So you might get used to being the jack of all
  318. trades and feeling very good about knowing, yeah, I could do this whole thing by myself. But when
  319. you find yourself in a collaborative environment, you often are and probably shouldn't be the best
  320. designer, the best developer, the best data journalist, all these things at once. So just
  321. learning to be comfortable and then kind of put your ego aside. Let people who shine at something
  322. let them shine at it. And not only will that make a better, a smoother and probably better
  323. process and probably better product in the end, it also will allow you to grow and achieve your
  324. preferred specialty. Something I might add to that is also from experience working here at Axios
  325. is that it helps to let the story steer the ship. As in every piece, every person who's
  326. contributing to the story is supporting the whole. And so the visual has to work with the words and
  327. has to work with the style and has to work with the social strategy. And all of those pieces are
  328. ultimately like our mission is to produce the best story possible and to tell a specific
  329. narrative. And it is humbling and also helpful to have that to fall back on when we're all trying
  330. to do our own little pieces of a story. Great. All right. Well, I think maybe the
  331. final question we can talk about before we go on to some Q&A section is I'd love if anybody wants
  332. to talk about how either technology or tooling for making visualizations has changed over the
  333. course of your careers and some sort of how has that impacted your work or maybe even what are
  334. you excited about for the future? Anything on the horizon that you're interested in in that space?
  335. I'm getting old. That's kind of scary to reflect on. I mean, I would say there's a few things. I
  336. think it's easy to take the cloud computing revolution for granted, just the fact that
  337. we're all able to just publish stuff from our MacBook with whatever flavor of platform you prefer.
  338. That wasn't the case when I started my career. You'd show up and they'd be like, you need an edge
  339. server in the basement to do anything and that costs X thousand dollars or whatever. That really
  340. I think opened up more than we appreciate, I think, in terms of opportunity. And then I just have to
  341. say, I think open source software just continues to be the seedbed and fuel or whatever the right
  342. metaphor is for all of it. I think sadly we've seen a decline of that in recent years in the
  343. news industry, which I presented a report on recently with my colleague Scott Klein.
  344. And I'm sorry to see that, but I'm proud to be here with Observable, who I know is one of the
  345. great leaders in that. So we just all need to contribute and keep investing in it because it
  346. does pay off. I think one thing that has definitely changed, I've been doing this for
  347. about a decade now, and the barrier to entry on getting to just getting started with a project
  348. is certainly lower. A lot of that comes from the open source stuff that was just mentioned. A lot
  349. of that comes from just more tools like DataRapper, like Flourish, like Observable, like all these ways
  350. that you can dive into DataViz or data storytelling or coding in general. But also just the amount
  351. of information there's. I mean, a lot of us I'm sure have had the problem of scouring the internet
  352. for some obscure Stack Overflow reference that had to fix some bug. And whether you're using
  353. something like whether using AI, whether you're using just Google getting better at searching,
  354. whether there's just more stuff out there, or maybe you're using like a Slack community.
  355. I think that it's never been a better time to start learning this stuff because you can get
  356. started so much faster. You don't need to struggle alone. There really is a very large community
  357. and ecosystem of information out there. Yeah, I think related to that is the amount of tools
  358. that we can keep on our tool belt is a lot bigger now because they have been so more established
  359. and there's a lower barrier to entry. And I think that's also exciting in the way that we can get
  360. more experimental with our data storytelling and make things that we imagine more real quicker.
  361. Yeah, totally. Great answers. All right. Time to move on to Q&A section. So we've got lots of great
  362. questions. I think a couple of these questions may have already been answered in the chat. So I'm
  363. going to jump around a little bit. So let me start with a question from Pedro. So Pedro asks
  364. about impacting the reader. In visualization, we generally encode information through visual
  365. variables such as symbols, shapes, sizes, colors, et cetera. This turns data about people into
  366. visual abstractions. How do you think we can humanize visualization in humanitarian content
  367. themes such as war, urban violence, discrimination, poverty, et cetera?
  368. I mean, I can just start. I think this is a huge and constant challenge that we face.
  369. I remember working on a project with some of their Axios folks on when we hit the million
  370. COVID deaths milestone. And I feel like that's also when this challenge came to the forefront
  371. for a lot of publications. How do we visualize, how do we get people to appreciate the size of
  372. one million deaths? And that might be a good example to look back on as different newsrooms,
  373. each one took their own approach. But I think it has to come from a place of empathy and trying to
  374. give people comparisons and ways in so they can relate to what might be colder numbers
  375. into something that they can understand and relate to. And then, you know, the impact happens.
  376. Yeah. I was just going to say that I think this is a question that's very important and it's hard
  377. because at the end of the day, as communicators, sometimes the most clear way to communicate
  378. something can come across as cold. And I think that can also, I think we think about this in terms
  379. of like databases, but the same thing can definitely be true in the traditional written
  380. journal or in video. And like, I think we all know examples of news content, whether it's a video,
  381. whether it's an article that just rubs you the wrong way or clearly doesn't have empathy or is
  382. clearly talking about the subject in a way that just misses the mark. And a lot of times it's less
  383. to do with what symbols are used or what colors are used and just how the general feel of it.
  384. So getting your content in front of people and sharing it with them. If you have a diverse group
  385. of voices in your newsroom, have them look at it, have them look at your content and see this
  386. anything strike you in a way that maybe we've missed. And I think that's probably a better
  387. approach than just saying, oh, here's our color palette that we think is good for this topic. Well,
  388. in all circumstances, is that going to be enough? Yeah, and this is a really great question and a
  389. big challenge. I know it's one that my colleagues in our excellent Reuters graphics department take
  390. super seriously. You know, I just posted in the chat a link to their work overall, but to two,
  391. I think, kind of outstanding examples of dealing with this. One by my colleague, Mario Zafra,
  392. and another by Ali Levine. And, you know, in both cases, you know, in almost all cases,
  393. our Reuters team in graphics is trying to use illustration and other more creative
  394. approaches to design, to humanize and to connect. And I think just showing people in Mario's case
  395. is one good example. You know what I mean? Not letting it be dots, you know, is one classic
  396. approach that goes back to the Isotype, you know, and some of the foundational tools of data
  397. visualization. And it's not my forte, personally. And so every time they put something out, I'm
  398. always kind of impressed and jealous of how they're able to bring their creativity and their artistic
  399. skills to bear. And I think we should, you know, as we can do more and more with data, and I can
  400. say more and more about cloud computing and Amazon, you know, lambdas and AI, blah, blah, blah.
  401. You know, we can't, we need more of that in journalism, but we cannot lose that artistic
  402. side of graphics, which is just so important. And, you know, I will not be replicated by AI anytime
  403. soon. Yeah, very true. Well, that's a great segue, maybe, into the next question. I knew it was
  404. coming. Okay. Generative AI. So, Ernesto asks, besides direct help in faster slash better coding,
  405. how do you think that generative AI is transforming the way we do storytelling with data?
  406. And it would be great if you have any examples.
  407. I think it's going to increase, I mean, to use kind of a loaded term, I think it's going to
  408. increase the accessibility of the type of work we do. It's going to make, it's going to bring
  409. more people into the conversation and allow and bring down the barriers or entry to just like
  410. figuring out how to get something on the internet or turn the data into a chart. And I'm seeing that
  411. every day here in our Reuters newsroom, where we recently launched a internal chatbot that through
  412. an MCP server is able to make charts and data wrapper just by having a chat. And we're going
  413. to pull in access to our financial data terminal next. So, you could have a chat that says,
  414. hey, can you get me the Apple stock price over the last 12 months and do it as a percentage change
  415. versus the S&P and Michael to make a line chart. And just by doing those kind of chatting maneuvers,
  416. you're able to publish a chart. We actually did one that way and put it on the homepage just two
  417. days ago. It was the first example. That one was by me because it's my job to be a full-time nerd.
  418. But I just see so much evidence that those types of interfaces or whoever you were going to call
  419. them are just going to put data and graphics tools in the hands of so many more people.
  420. And there'll be bad things to come out of that, but I suspect there'll be a lot more good things.
  421. Yeah, I think like obviously there's so much that can be said about how the news
  422. can use AI. But thinking for me personally, I tend to approach it just as like another tool,
  423. and every tool has its limitations. Every tool is good at something and bad at others. And what I
  424. would hope is that you will see more data storytellers using it for the things that humans are
  425. bad at and maybe not letting it just replace the things that hopefully everybody in this call like
  426. the fun parts of the job, the interesting parts of the job, and just trying to find that balance
  427. for you as you tell stories. Can this improve your work? If so, that's great, but just be aware
  428. that it is a tool and tools require the user to know what they're doing. Yeah, I recently had a
  429. conversation with Marco Hernandez from the New York Times, a great graphics artist at the School
  430. of Visual Arts. And he said, let the AI do the monkey tasks is the way he put it. But you always
  431. have to remember that you are the protagonist, which I thought was quite a zinger from Marco.
  432. Yeah, I think one way we're experimenting with AI at Axios is one of our reporters set up a chat bot
  433. that we fed in kind of some standards around data visualization and data that we have. And
  434. it's basically to help reporters triage a data set when they encounter it in the wild to see,
  435. you know, is this suitable for visualization? Is there any like pitfalls that might, you know,
  436. like missing data or things to be, you know, to be aware of when considering whether to cover it?
  437. But it's far from being gritty for prime time. But it's an interesting idea. I think
  438. it's exposing kind of how nuanced our job is, because there aren't always set rules for,
  439. for, you know, evaluating data sets. But even getting a reporter like 80% of the way might be
  440. helpful too. So worth both sides of worth pursuing, I think. Totally. All right, I know we're going
  441. on a little long here. But I do just want to get to this final question that we can close out with,
  442. because I know this is a question I get a lot of questions about this personally. And so I know a
  443. lot of people are curious about behind the scenes, like kind of a little bit of the plumbing of of
  444. how these visual stories and data journalism comes together in the real world. So there's a
  445. lot of aspects to this question. I think we can all go around maybe and talk about one or two of
  446. them. So current asks, I'm curious about tooling and collaboration in newsrooms. Does it look like
  447. Figma designers and then devs implementing those mocks? Or is it more like designing in code?
  448. How do you share work in progress? Do you have an internal staging environment?
  449. How do you share multiple prototypes internally for feedback? How does it migrate to production?
  450. There's not going to be one answer. You know, newsrooms are kind of all over the place.
  451. Everybody is is solving the same problems in slightly different ways. And so we could probably,
  452. if we tried with some post-it notes, come up with like, here's the 25 tricks, and then everybody,
  453. every newsroom is some mix and match of those 25 would be my hunch. In my newsroom experience,
  454. which is now more than 20 years of doing this, I've yet to encounter a newsroom team that runs
  455. like a truly formal product team, you know what I mean with that type of division of labor. It's
  456. more and more common for a team like ours to partner with our product team. And then they
  457. kind of force us into their their agile process or, you know, whatever you might call it. I think
  458. there's a lot of merit to that. But I think it's pretty unusual to see that in the newsroom, which
  459. which tends to run in a little more of a chaotic fashion. In my experience, and just having done
  460. this about a lot of people, I think you're going to see a lot of website development that's story
  461. specific being done in a static site framework that is roughly like observable framework,
  462. to be honest with you, it's going to be some strange Frankenstein combination of node and
  463. templating tools, maybe involving svelter react, maybe still being on an old templating tag thing.
  464. And every newsroom has its own odd variation on that that is then used to bake out or, you know,
  465. build a static distribution. That's just such a sturdy approach to publishing custom
  466. micro sites or whatever you want to call them, that I think it's pretty much been irresistible,
  467. you know. But there are many other ways things happen. And I don't want to
  468. bore you by doing a laundry list. But that's my impression. What do you guys think?
  469. Yeah, I think that it's always going to depend on the project and the person building it,
  470. specifically looking at the design by code or do it in Figma. I think every single time,
  471. especially when it's like a data heavy story, the data very much will inform the shape of it.
  472. If you do the whole thing in Figma first, whenever I'm developing, I'm thinking, man,
  473. I really wish we knew more about the data before we made these Figma designs. And then the other
  474. way around, when I'm designing my code, it's like, wow, this would be a lot easier if we had
  475. some cleaner prototypes designed out. So I think it just really is going to be on a case by case
  476. basis. Because especially if you're not doing the same thing every time, and it is more bespoke,
  477. more bespoke visualizations, more bespoke data story experiences, there's always going to be
  478. nuances that you probably wouldn't find in traditional static site building.
  479. Yeah, I think our approach varies based on timeline and who's leading the project. I think
  480. everyone is familiar with a slightly different suite of tools. So what I might do in a spreadsheet,
  481. my colleague might do an R, but we end up in the same place. I think in terms of feedback,
  482. Axios is kind of unique nowadays in that we're fully remote and a pretty small newsroom
  483. comparatively. So everything we do is on Slack, for better or worse. And a lot of the processes
  484. we have are kind of cobbled together slash what we personally think is the best approach at the
  485. moment. But that's also playing to our strengths. It's so unique to us as a team. So I think that's
  486. yeah, just echoing what folks are saying, it's going to vary.
  487. Yeah, totally. All right. Well, thank you so much to everybody on our panel here today. I think it
  488. was a fantastic discussion. Thanks to everybody for watching, for submitting your questions. We are
  489. sorry if we didn't get to your question. But just a reminder, the recording will be sent out after
  490. this. So no need to worry about that. And yeah, once again, thanks, everybody. And we hope you
  491. can join us again in the future for another discussion like this. All right. Bye.

Downloads

Recording video · Timestamped transcript