Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

piratebroadcast 131 points 69 comments October 09, 2026
jessewaites.com · View on Hacker News

Discussion Highlights (11 comments)

whythismatters

>To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations. https://github.com/jessewaites/antiquity

acgourley

Very cool. I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing. Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.

jvanderbot

IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.

dang

Recent and related (by HN's own https://news.ycombinator.com/user?id=benbreen !) Using Opus 5.5 to discover a new eyewitness record of the dodo - https://news.ycombinator.com/item?id=49926917 - Oct 2026 (79 comments)

sorokod

If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run. Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...

jttnr

That was a fascinating read, I really enjoyed it. Literally like exploring lost knowledge. Great work and a great write-up. I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates). I wonder what else could be found in such archives. Some ideas: - Locations or routes of sunken ships and their missing cargo? - Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain? - Unusual weather events, like snow in the summer? (edit: formatting)

yieldcrv

I love this, one major friction I’ve seen to human coordination and advancement has been the journals in different languages Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all. Articles that are stubs on the English page may be massive fully fleshed out articles in another language, and nothing native to the site or anything I’ve seen will tell you that there is more information in one variant LLM’s can find the word associations and compare them in all languages, even if it itself doesn't innately know language and there would be so much low hanging fruit here like this engineer found

fudgybiscuits

Sorry he found an unrecorded volcanic eruption in some records? He found "A new eyewitness record of the extinct dodo" that's 400 years old? These seem to me to be rather dubious claims.

wavewrangler

In my experience, when a disruptive tech comes along that displaces a certain way of existing, all of the former examples of this happening have resulted in people adapting. That's what must happen. And if you think LLM's are lessening the value of previously valuable work, then it is time that you increased your own output to once again be high and above that which LLM's are replacing or to what you perceive as having lost value. You can do that, that is totally within you, but you have to find it for yourself. Or you can just continue talking smack and contributing to nothing. But then your output is going the opposite direction of what you say LLM's are taking away to begin with. These are conflicted times we live in, but they really don't have to be. For the record, I think this project was an excellent presentation, I loved the interactive elements snd the meteorite (or is that a meteorwrong?) to come across the page. This is about as good of a use of LLM's as I have seen. I didn't see anything about it in the article, but does anyone know what future plans are for this particular project, or is that a wrap? I didn't see I the Where This Stands section anything about future search topics. This really feels like a time where finding the question is every bit as important as finding an answer to that question.

Ariarule

Wonderful work, and it's unfortunate that knee-jerk anti-AI sentiment is getting in the way of appreciation for the accomplishment by some others. Consider this: if the exact same oddities and anomalies had been discovered 3 to 5 years ago with traditional NLP, or OCR and/or some clever statistical techniques, would anyone be dismissing it? The post goes into the work they did to orchestrate and explore with the AI in this case, so it's no less impressive an effort than the hypothetical, and the discoveries just as valid for consideration.

jamienk

This article is NOT like the Dodo or Newton ones. The Dodo author was an expert in Dodos. The Newton one was by a Newton expert. In those cases, it was the relentlessness of the AI that found edge cases that were interesting. But Here, the start is "What field should I approach and what questions should I ask?" But this is NOT the same! When I watch my non-tech friends vibe-coding, I'm struck by how they are so ignorant of basic tech stuff, but little tiny bubbles of "experience" start to percolate in them after a while. Things like "wait, I think this project has weird dependencies that are going to cause trouble when I move it to my office machine" or "hold on, the mobile version doesn't share the same text with the desktop version?" It is much less interesting to have an attitude of "I don't know or care, just give me a 'result'" vs "Aha! I see where we could apply this in a productive way!" This is like the Anthropic people feeling like they found many many "very important" Linux bugs https://www.youtube.com/watch?v=NnV_cWeoo5Q - hint: no. When it's something you know about, the overstating is obvious! It's only when you don't know about things that shallow work (or slop) feels significant.

Semantic search powered by Rivestack pgvector
8,999 stories · 84,560 chunks indexed