Introducing Storytracker

By • • Dodging the Memory Hole in Columbia, Missouri

Slides

Show the extracted slide text

Slide 1

StoryTracker 1.0
A brief introduction by Ben Welsh

Slide 2

My name is Ben
Sometimes I go by @palewire

Slide 3

I work here
The @latimes in #DTLA

Slide 6

TK

Slide 12

Screenshots
 are good

Slide 13

Screenshots
 are good

    But
also stupid

Slide 14

Screenshots
 are good

     But
 also stupid

   HTML is
the gold mine

Slide 17

>>> import storytracker
>>> obj = storytracker.archive('http://www.stltoday.com', output_dir='/home/ben/Desktop/')
>>> obj.gzip_archive_path
'/home/ben/Desktop/http!www.stltoday.com!|!!!@2014-11-08T23:49:08.188549+00:00.gz'

Slide 18

>>> obj.write_hyperlinks_csv_to_path('/home/ben/Desktop/stltoday.csv')

Slide 19

>>> obj.write_overlay_to_directory('/home/ben/Desktop/')
'/home/ben/Desktop/overlay-http!www.cnn.com!|!!!@2014-11-08T23:49:08.188549+00:00.png'

Slide 20

>>> urlset.write_overlay_animation_to_directory('/home/ben/Desktop/')
'/home/ben/Desktop/urlset-overlay.gif'

Slide 21

>>> urlset.write_href_overlay_animation_to_directory(
>>>     'http://www.stltoday.com/sports/columns/bernie-miklasz/bernie-big-red-thrive-under-you
>>>     '/home/ben/Desktop/'
>>> )
'/home/ben/Desktop/urlset-href-overlay.gif'

Slide 22

So what?

Slide 23

So what?


Automated
 analysis

Slide 24

So what?


Automated
 analysis


Here’s how

Slide 26

Who does Drudge love the most?

Slide 27

Who does Drudge love the most?

Nov. 4

Slide 28

Who does Drudge love the most?

Nov. 4      7 days

Slide 29

Who does Drudge love the most?

Nov. 4      7 days     96 archives

Slide 30

Who does Drudge love the most?

 Nov. 4     7 days     96 archives


953 links

Slide 31

Who does Drudge love the most?

 Nov. 4       7 days      96 archives


953 links   448 stories

Slide 32

Who does Drudge love the most?

 Nov. 4       7 days      96 archives


953 links   448 stories   142 domains

Slide 33

9.

Slide 34

9.   8.

Slide 35

9.   8.   5.




5.   5.

Slide 36

9.   8.   5.




5.   5.   4.

Slide 37

9.   8.   5.




5.   5.   4.




3.

Slide 38

9.   8.   5.




5.   5.   4.




3.   1.   1.

Slide 41

>>> import storytracker
>>> obj = storytracker.open_wayback_machine_url(
>>>     'https://web.archive.org/web/20010911223655/http://www.cnn.com/'
>>> )

Slide 42

What next?

Slide 43

What next?

Real world
 feedback

Slide 44

What next?

Real world     Grow, via
 feedback     partnerships

Slide 45

What next?

Real world     Grow, via      PastPages
 feedback     partnerships   integration

Slide 46

What next?

Real world      Grow, via      PastPages
 feedback      partnerships   integration

SVG outputs

Slide 47

What next?

Real world      Grow, via      PastPages
 feedback      partnerships   integration

SVG outputs    Storysniffer
                 refactor

Slide 48

What next?

Real world      Grow, via      PastPages
 feedback      partnerships   integration

SVG outputs    Storysniffer      Your
                 refactor     wild ideas

Slide 49

http://storytracker.p
astpages.org




                        Questions?

Slide 50

Sorry, haters.
Still standing.

Recording

Show the timestamped transcript
  1. Hi guys, hey Ted. So this has been a fun little trip for me to try to
  2. parachuting into your world and I really had a lot of fun listening and I just
  3. I'm gonna take a few minutes to walk you through something kind of weird in mind
  4. all right we're gonna call it story tracker what is it well first who the
  5. heck am I my name is Ben Welsh I sometimes go by pale wire online I'm
  6. originally from Eastern Iowa and at one point in my life I was a student here at
  7. Mizzou and I've had a real fun time the last couple days coming back to check in
  8. on everything I'm pleased to see Jesse Hall still there the J school's gotten
  9. bigger there's a thing now called alley a that I've learned about which is
  10. pretty exciting and it's been I'm now discovered I've turned into that old
  11. crank who leads his wife around his college campus and it's like you know
  12. things just ain't what they used to be and that's kind of said it's been fun
  13. for me the last couple days to grow in that way and also my day job is in Los
  14. Angeles where I work at the LA Times and I'm one of those nerds in the newsroom
  15. who does a little bit of computer programming a little bit of writing and
  16. tries to take data and flip it around and do stuff with it to make news writer
  17. to make stories happen to make web applications happen to make graphics to
  18. do stuff with data right and so that means that I have to I do a lot of
  19. computer programming which I started to learn here at Mizzou thanks to irie
  20. night car and their program there which I would really recommend and one thing
  21. I've always loved about newspapers since I was a really little kid is big blowout
  22. front pages right who doesn't love them and here's one of the most famous ones
  23. in LA Times history and when you talk about archiving this story you know in
  24. the database that might go into in the metadata and all the things that we're
  25. discussing at this conference right you can easily imagine what would go into
  26. the database there's a headline there's a pub date there's a byline there's the
  27. big blob of text that it is everything that went into the story but one thing
  28. that sometimes isn't captured or that is sort of the the intangible thing or the
  29. thing that isn't always in the metadata that's really powerful about this front
  30. page is the design right it's the layout it's the big block headline it's the
  31. context of knowing that's not normally what the front page looked like which is
  32. sort of an unwritten part of the design that's in there as well and that's
  33. something that's just special about newspapers right and that's why we love
  34. these big front pages like that and you know I work on the website and I'm a
  35. computer nerd and I think sometimes that's often true about news home pages
  36. in the same way that in the home page is embedded a design that is part of the
  37. message right in that layout of how the page is done and I really started
  38. thinking when the Egypt thing was happening a couple years ago and
  39. everybody was doing these big crazy home pages to try to communicate to
  40. Americans what a big deal this thing was that was going on over in Egypt that
  41. wow I don't think that part of news presentation online is necessarily all
  42. that well captured and I started really thinking about it and I took a bunch of
  43. screenshots myself just by hand with a little drag-and-drop just for fun and I
  44. started thinking back to another event which is one of my favorite news home
  45. pages from a couple years before you guys remember balloon boy this was the
  46. hoax news story in Colorado about Falcon had supposedly lifted off air I was
  47. having a bad day at work that day you know I was kind of down and then this
  48. happened and it really rose you know brought my spirits back up and so I'd
  49. taken this screenshot at the time but what then when the Egypt thing happened
  50. and I was really kind of thinking about are we saving this stuff I went to the
  51. Internet Archive which is one of my favorite sites in which I direct every
  52. investigative reporter I work with to to dig stuff up all the time I said I
  53. wonder if they captured balloon boy because it was such a great moment for
  54. me and it wasn't there and you know it's not any fault of theirs they have this
  55. big broad mission the entire goddamn internet you know what I mean they can't
  56. get everything and to catch every little thing especially the fun screw-ups on a
  57. news home page you'd have to be hitting that thing all the time right because
  58. these things are changing more and more and if you any of you working in
  59. newspaper website you know that there's pressure to get that thing just
  60. churning because you want people who come back to feel like they're getting
  61. something fresh you want to trick the search engines into thinking you got you
  62. know fancy fresh news by putting it there because they waited and there's a
  63. whole game that goes on in newsrooms to keep that thing moving I just started
  64. thinking I wonder if we could watch that or track that and since I'm a computer
  65. nerd I decided to try to write some code to do it and I made a website the
  66. websites called past pages org and Edward explained it a little every hour
  67. it goes to about a hundred news home pages all around the world and it just
  68. takes a picture all it does is the server which you know just has a little
  69. web browser inside of it and it fires it up it waits for it to load it takes a
  70. picture and it saves it I said so it can be done even I can do it right in my own
  71. crappy code in my crappy website which I then put all up online and that was
  72. kind of fun that was like a fun weekend I had hacked this thing that did
  73. something that was cool but then the bill came well how am I gonna pay for
  74. this or even should I pay for it and is it worth saving or keeping does anybody
  75. want it like I don't know but I knew I wanted to make the point that we should
  76. be saving these home pages because they carry really important editorial
  77. decisions they reflect the hierarchy and the choices made by news gatekeepers
  78. and I didn't think they were getting saved often enough so I did a
  79. Kickstarter and I really did not think I'd raise any money right I just figured
  80. this is gonna be so I can tweet it out and make a bunch of noise and be like
  81. we should be doing this right and then I was I got shocked to actually raise the
  82. money so I had to figure out what the hell I was gonna do right and so so the
  83. money paid to keep the website up for a little while longer and I tried to start
  84. talking to different people and figuring out well is this something we
  85. should do or not do or what what have I stumbled into and and in that course I
  86. ended up talking to some people here at the Reynolds Journalism Institute Randy
  87. and Edward just about like hey what do you what do you guys think of this you
  88. have this new archiving mission this fit what you want to do and they were like
  89. yeah maybe sure let's try something right you know let's let's yeah and so
  90. what we tried is what I'm here today to unveil which is a new what well in these
  91. backup sorry so the screenshots are good right so when you take a picture when
  92. this is from past pages it looks good right but it's just like a flat stupid
  93. image right it's like a you know and you can't the data inside of it is just all
  94. kind of flat you can't really parse it out very well right but of course inside
  95. the web page is the HTML right which contains all of the layout and code
  96. there's that big headline that big headline right there is that h1 tag
  97. right there right and there's the style files that kind of do all that layout
  98. and make it happen and we you know we quickly realize in our conversations if
  99. we want to get if we really want to analyze what the choices are what's
  100. going on we got to get beyond the image we got to be able to get into that data
  101. that's the HTML like they do at the Internet Archive and start to analyze
  102. that and so we came up with this new idea story tracker which is really kind
  103. of an experiment and it's hitting sort of version 1.0 today where I'm going to
  104. show you some of the features we're going to talk about what it does what
  105. we're trying to do and then you can tell me what sucks about it and how we can
  106. make it better okay so that's going to be the rest of my spiel so as of today
  107. you can go to your terminal and there's going to be code guys just a little and
  108. I think I'm the first one to show code which I always love I'm really proud of
  109. that so if you were to go to your computer terminal this is going to be in
  110. the Python programming language which is not as scary as it sounds and you could
  111. just with one command install our friend story tracker right and then once
  112. you have an install it's just a simple library that takes some of these tasks
  113. and things we want to do and tries to make it easy so the first thing it does
  114. is just a really simple archival system you know before we can analyze the HTML
  115. on the page we got to download it we got to save it we got to put it in the
  116. place we can do so with one little command like this you can download any
  117. page not even a home page really any page right it'll archive it in the
  118. structured way zip it up and say so st. Louis post-dispatch right this is going
  119. to be our example from a couple days ago this was their home page and then it has
  120. code it's then going to fire up HTML inside the system and extract all of the
  121. hyperlinks that appear on the page so you're able to take out of any of the
  122. home page zip out all the hyperlinks and dump those to a spreadsheet but it
  123. doesn't just take the hyperlinks out which really any parser can do that's
  124. really pretty elementary parsing thing like here's the URL here's the headline
  125. that was inside of it right that's pretty easy to do but what we're also
  126. doing is the code is leveraging some browser simulators to actually fire up
  127. and render the page just like the image does right and pull out where it is on
  128. the page so where did this link appear how big was the font what box is around
  129. right is it does it jump to a story or does it jump to a section front or an
  130. advertisement or something the system estimates not to be a story so in
  131. addition to parsing it it's going to pull out all this other metadata about
  132. the layout on the page the treatment and prominence it's given and then whether
  133. it's not not it's a story and then that gives you then structured data that
  134. will allow you to analyze how the page was laid out in a way that's a little
  135. more sophisticated right and so that is kind of the intel inside of the code
  136. that we've made and then what's coming next is the sort of some basic examples
  137. of what you can start doing once you once you get there so the first thing
  138. you could do is you can highlight where the stories are and stories aren't right
  139. and you can overlay on top of a page or really anywhere you wanted where these
  140. stories are and that's what we kind of pulled out and it just will create an
  141. image like this for you on the fly and then you can take multiple screenshots
  142. and start to like automatically make an animated gif just like of cats or of
  143. anything else on the internet right that's showing those changes over time
  144. on the page so one line of code boom you got it right then you can say even track
  145. an individual story over time and see how it moves down across the page and
  146. that's not just something you can spit out as an image you can then also extract
  147. that data and automate analysis of the quantitative data that comes with it
  148. right in addition to visualizing it and that is what really gets me excited it's
  149. you know beyond making a goofy little gift you can start to run this at really
  150. large databases of home pages and ask them interesting questions so I did a
  151. little case study right and our case study today of what you could do with
  152. something like this is one of my favorite websites of course the drudge
  153. report is everybody know the drug report I think you probably do the key thing to
  154. know about it for this example is that it doesn't write very much original news
  155. it's mostly links to other third-party sources but it's such a popular aggregator
  156. homepage that many many people come here and click on it so the selections they
  157. make has a large influence on media stories to get attention right and that
  158. leads me to and so and also past pages archives that every hour not just the
  159. screenshot but now thanks to story tracker for several months it's also
  160. been archiving the HTML and so we can lean on that database that we've
  161. created of archives things over the last couple months and ask and answer a
  162. question right and this is what's gets me it gets me excited because when I was a
  163. graduate student at Mizzou I had to do a content analysis right to earn my
  164. master's degree and that involves studying newspaper coverage of political
  165. candidates and I was so boneheaded at the time I literally went through and
  166. counted them like by hand and it took me a week to do all the bean counting
  167. necessary to run a story to answer a question like who did Matt drudge say
  168. linked to the most and and but when you have that all automated and you're able
  169. to quickly run through it you can answer that question I literally did this case
  170. study in about 30 minutes on Saturday afternoon because the date is there the
  171. codes there that's able to run through and do it she just kind of piece it
  172. together and you slap it out so November 4th was the election right the midterm
  173. election big day for Matt drudge right our study is going to take the previous
  174. seven days of home pages of the drudge report for that period and in there
  175. there's 97 archives that were taken down by the past pages system I fed those in
  176. the story tracker had it extract all of the unique URLs that ever appeared in a
  177. home page because it doesn't necessarily change all that much from hour to hour
  178. but over time you're going to have quite a bit of turnover of the links so 953
  179. distinct links appeared on the page in that period now if you use the drudge
  180. report you know that the bottom half of the page is just a bunch of filler right
  181. it's the same links to the same people that never change doing an analysis you
  182. so for do any meaningful analysis you have to be able to subtract out which
  183. ones are stories in which one aren't guess what story tracker does that for
  184. you right with the estimation we do based on the URL of what's not a story
  185. so of those 953 links 448 are estimated to actually go to new stories distinct
  186. links right and of those there's one hundred and forty two different news
  187. domains so over a seven-day period prior to the election
  188. Matt drudge linked to about a hundred and forty two different news sources but
  189. who got the most links any guesses
  190. okay Washington Post is one guess Fox News okay CNN okay AP I love this this
  191. is like my alright well let's see so let's do the countdown number ten
  192. number nine USA Today right followed by the hill.com believe it or not which
  193. does a lot of coverage of what's happening in Washington there's a tie
  194. there's a followed by a three-way tie between Politico and two British
  195. tabloid newspapers right maybe surprising number four the New York Times
  196. didn't do so hot number three is the Washington Times right-wing paper in DC
  197. and then we actually have a tie for first place between the Washington Post
  198. and believe it or not Breitbart.com which people who if you guys know what
  199. that is Breitbart.com was a new startup founded by a former person who worked at
  200. the drudge report and they have connections and so they get lots of
  201. links which is something you don't you can't really prove or know until you
  202. quantify it in this way so this is not some amazing awesome study that proves
  203. anything but it just shows that when you're able to automate your analysis
  204. of those types of choices you can really quickly start to ask and answer
  205. interesting questions you know what I mean you could you could also ask hey
  206. show me every photo of Michelle Obama Matt Trudges ever run might be a thing
  207. that you could also extract right which I think would be a very interesting
  208. photo gallery or you could compare how often different entities or terms
  209. appeared on different news home pages of different outlets you could say how was
  210. Ferguson covered by certain newspapers or broadcast websites compared to print
  211. ones you could start to really kind of put together a more complex analysis of
  212. what type of media coverage there is right and so this drudge analysis I did
  213. was again about a half-hour work the codes also up on github along with the
  214. rest of the past pages stuff and story tracker if you're interested and back
  215. again to our friends the way back machine so let's say oh my gosh Ben if
  216. I'm going to use this that means I have to be archiving all this stuff right
  217. that's a big bunch of work it's only going to be going forward it's no good
  218. but the way back machine has tons and tons of stuff and with a simple one line
  219. of code you can suck in any URL from the way back machine stick it into story
  220. tracker it will this image came right out of story tracker just then creates
  221. it on the fly and you could it could also then parse out all the URLs and do
  222. the same type of analysis over a lot of older sites so you could say to look
  223. back over time and say how has the design of news home page has changed
  224. there are more links less links do they move faster move slower what type of
  225. things do they cover or don't cover could very easily be done by kind of
  226. taking this and pointing at other archives and that type of thing could
  227. potentially be integrated with other sources as well I just wanted to do this
  228. as a demonstration so you know what's next this is just some basic open source
  229. code it's an experiment that we're doing together to try to create something that
  230. doesn't exist and have good is you know that doesn't exist today but it needs
  231. feedback from people like you who know this stuff and want it to be in want
  232. maybe want or don't want things like this to let me know how we can improve
  233. it if we're going to make these efforts and I think that that's the that's the
  234. approach I'd like to take is how can we grow this type of effort we're making
  235. with our GI to make code that does stuff that isn't currently done and
  236. integrate it with different things in the world and the only way that's going
  237. to be useful is if we find people to link up with you know I mean because I
  238. was really inspired by the media criticism work that I did in graduate
  239. school and that sort of thing that's kind of the natural partnerships I've
  240. had in mind but I'm open to anything I'm here just to kind of talk about it you
  241. know I'm really going to that my next push is to try to integrate this further
  242. into past pages so why can't we be automating the type of analysis we talk
  243. about there so that it happens on the fly on a rolling basis in different
  244. interesting ways and I've got a few ideas how that's going to happen in the
  245. next couple months we could we don't have to just output images we could
  246. output SPG files that would allow more interesting animations or people to do
  247. different stuff the system that detects what story or what isn't a story we've
  248. extracted from the library is actually an independent library that could be
  249. used for a lot of different things and I think there's some really cool machine
  250. learning kind of things we could do to take that to another level and then any
  251. other wild ideas that I hear anybody has I'm totally open to and I think would be
  252. cool so that's it that's story tracker it's online today you can go there
  253. there's full documentation of how to use it as a programmer including numerous
  254. examples of it in use and yeah that's it that's your spiel so you can feel
  255. free to tell me it sucks or what you think it's useless or how it could be
  256. more useful or what you just think I'd love to hear it okay we cannot use the
  257. mic thanks it looks great I wonder if so at the Library of Congress that we
  258. collect like election websites from candidates and you know back to in the
  259. late 90s so I wonder if this could if if the if this could be useful for
  260. analyzing things in those archives not that there's really story elements there
  261. but you know there are different other elements well the way it's built it's
  262. it's agnostic to what type of pages being put in you know I mean just all
  263. the examples and the way I've kind of framed the code have been built around
  264. home pages but the code doesn't care you could feed it any set of pages and it
  265. could do the exact same stuff two questions have you thought about applying
  266. this I know that reporters always ask like how do I like figure out a way to
  267. track whether this campaign website has changed over what period of time have
  268. you thought about applying this to that and the other question was do you have
  269. archives going back to what was going on during Ferguson because I was really
  270. curious about the way that the news broke particularly over like st. Louis
  271. today or any of those sorts of things since there's this like meme that
  272. Twitter is much better at covering like this early breaking news and I actually
  273. thought that the media in st. Louis actually did a pretty good job once
  274. this right became national news sure so the campaign question the first one I
  275. actually hadn't really thought about it so you guys are asking it now I think
  276. it's totally plausible and do it you just would need to set up the system to
  277. begin the archiving or you need to plug in somebody else's archive which would
  278. totally be doable I just hadn't thought about it it would work the Ferguson
  279. thing past pages now has 1.7 million screenshots which is kind of scary and
  280. and it includes the whole Ferguson period the images I didn't have the post
  281. dispatch I have an archive that which is just an oversight of selecting the
  282. hundred sites because of the limited funds I have not expanded the number of
  283. sites that it archives a whole lot because everyone I add is just going to
  284. increase the the long-term cost but the screenshots for all the other major ones
  285. are there HTML archive HTML archival win a long line around that same time but
  286. that's only for a smaller subset of the sites but if you had pages or the things
  287. where they're useful you could very easily do that kind of analysis you know
  288. what I mean you could find you could identify certain terms right or a period
  289. of time and you could analyze what stories were given the most prominence
  290. what terms were used to describe it what photographs were chosen you could begin
  291. to start doing that type of work using this and I'd love to make it more and
  292. more useful for that kind of thing because that's that's that's the thing
  293. I'm really passionate about is how do we automate media criticism so we can ask
  294. bigger and more insightful questions yeah I mean just really quickly I think
  295. this is super cool because essentially like what Google started doing is not
  296. just indexing HTML but actually looking at layout and those sorts of things to
  297. seeing that you've built a framework to start looking at that is actually I
  298. think way more powerful than I had actually assumed that this was and if
  299. you're if people are into the gearhead stuff it's actually using a tool called
  300. Selenium so there's a really awesome software tool that is open source that I
  301. just took which is used to simulate websites for testing it so if you work at
  302. Facebook and you make one minor change you could break a hundred things on the
  303. website so there's this these crazy software systems that fire up the site
  304. and click all the buttons just to see what's broken and we're using that same
  305. software for like an entirely different purpose here of simulating the page and
  306. then pulling out the layout yeah and have you extended this to the mobile
  307. world as well I mean since you're getting the style sheets a lot more
  308. people are getting that info you know on their phone right it's totally possible
  309. so past pages only takes one screenshot per hour per site but the HTML when you
  310. open it up saying the tool Selenium I just described so I'm gonna have a fake
  311. browser open it up and render the page you can set how big the browser is when
  312. it opens so you theoretically could identify a set of viewports I guess
  313. would be the technical term right a set of sizes that you wanted the browser to
  314. open up at and render it and then compare those or just do one versus the
  315. other currently I just use the what supposedly the most common one 13
  316. whatever but you could you could take this into the exact same thing for any
  317. size any size that you wanted just by tweaking the code so one of the one of
  318. the things that I've been thinking about and I you know other people have
  319. thought about this too is when we're talking about archiving websites how
  320. often like if you're just gonna do a you know sort of straightforward stupid
  321. right capturing of it but you know in order to save space you kind of have to
  322. be strategic about that right at the same time if you have a breaking news
  323. story and things are being updated really quickly you don't want to I mean
  324. you kind of have to figure out I don't want to miss something really important
  325. right so how would we that's a research question I think you could use this
  326. and other tools to answer you know what I mean I don't know the answer but I
  327. think if you were harvesting the pages very frequently and they're just doing
  328. some basic comparisons of the changes you could pretty quickly identify what
  329. the rate of change is for different news sites and then use that to sort of one
  330. know that and that's it like an academic finding you could have but you could use
  331. that to try to find an optimal time to change it also you could there's ways
  332. you can do it with computers where you request the page you compare to the
  333. previous one if there's no change just don't save it you know what I mean it
  334. just kind of opts out of saving it I think there's also there's compression
  335. tricks that that I'm not using that I think could probably really reduce the
  336. size where like there's tricks for compressing files where you only compress
  337. the change from file to file you don't compress the entire file itself each
  338. time not to get like real geeky that that some people use and other people
  339. don't that reduces the amount of data necessary to save a series of files
  340. which would be a really cool thing to make easier to do yeah I'm not like that
  341. I'd be interested in this at all yes way to go so I have a suggestion question
  342. first I want to call you out to the rest of the people here this is the kind of
  343. person and the kind of interesting stuff that we love it's in that libraries
  344. another archivist should love and encourage and support and foster and do
  345. whatever we can it's the people who are essentially what we call the middleware
  346. people people who are creating access enhanced access points and doing
  347. something creative with it and more importantly he should opening the door
  348. for other people to take his code and do something creative with so the fact that
  349. you're here with a superb app is great but as a symbol for what many people in
  350. this audience want done with their their their archiving and their material
  351. that's your even better example so I just want to suggest another iteration
  352. not for you but anybody who contacts you to build off of there's increasing
  353. interest in the personalization of the web where right depending upon your
  354. cookie array right and geolocation by IP and other kinds of attributes you see
  355. a different web you see a different news you see different stories and you see
  356. different ads and as you know the Internet archives really interested in
  357. political ads as well as other ads and there's a growing interest it's
  358. certainly nascent at the moment to create essentially kind of a honeypot
  359. botnet where there are simulations of computers with different profiles that
  360. are constantly hitting up a variety of sites right to take these kinds of
  361. snapshots and take a look at assessing quantifying just how different is the
  362. web to what the site perceives you to be into the advert side advertisers perceive
  363. you to be I'm just offering that I think it's a great idea I think it would be a
  364. great research question I don't know the scholarly literature of whether people
  365. have you know done thorough studies of that or not but to me that's the kind of
  366. thing I'm excited about where if there's some partnership where somebody
  367. wants to do a study and they just need a little extra programming to kind of like
  368. make it happen you know what I mean we just need to stick a cookie on this
  369. thing and do a loop that the cookies not there right now you know what I mean if
  370. we could figure that out that's the type of thing that I would really love to do
  371. as like a next-gen or a spin-off of this so we have sort of a practical project
  372. that leads to sort of just enough web development to get it done you know and
  373. in the light of ethical considerations I have to raise the point that the
  374. speculation about doing that comes close to what's already being done in a class
  375. of click fraud where people are intentionally hitting up ads with
  376. different profiles in order to gain the revenue from those ads so like all
  377. interesting creative things there are with great power comes great
  378. responsibility okay we're gonna wrap up we're a little bit over

Downloads

Slides PDF · Recording video · Extracted slide text · Timestamped transcript