Hacker News SQLite | Not Hacker News!

Discussion (217 comments)

Showing 160 comments of 217

10 days ago

6 replies

Community, All the HN belong to you. When I made HN Made of Primes I realized I could probably do this offline sqlite/wasm thing with the whole GBs of archive. The whole dataset. So I tried it, and this is it. Have Hacker News on your device.

Go to this repo (https://github.com/DOSAYGO-STUDIO/HackerBook): you can download it. Big Query -> ETL -> npx serve docs - that's it. 20 years of HN arguments and beauty, can be yours forever. So they'll never die. Ever. It's the unkillable static archive of HN and it's your hands. That's my Year End gift to you all. Thank you for a wonderful year, have happy and wonderful 2026. make something of it.

carbocation

10 days ago

5 replies

That repo is throwing up a 404 for me.

Question - did you consider tradeoffs between duckdb (or other columnar stores) and SQLite?

keepamovin

10 days ago

4 replies

No, I just went straight to sqlite. What is duckdb?

simonw

10 days ago

1 reply

One interesting feature of DuckDB is that it can run queries against HTTP ranges of a static file hosted via HTTPS, and there's an official WebAssembly build of it that can do that same trick.

So you can dump e.g. all of Hacker News in a single multi-GB Parquet file somewhere and build a client-side JavaScript application that can run queries against that without having to fetch the whole thing.

keepamovin

10 days ago

In that case, then using duckdb might be even more performant than using what we’re doing here.

It would be an interesting experiment to add the duckdb hackend

1vuio0pswjnm7

10 days ago

1 reply

"What is duckdb?"

duckdb is a 54M dynamically-linked binary on amd64

sqlite3 is a 1.7M static binary

DuckDB is a 6yr-old project

SQLite is a 25yr-old project

duckdb reads Parquet or DuckDB

sqlite3 reads SQL

DuckDB is OLAP

SQLite is OLTP

1vuio0pswjnm7

10 days ago

I like SQLite

cess11

10 days ago

It is very similar to SQLite in that it can run in-process and store its data as a file.

It's different in that it is tailored to analytics, among other things storage is columnar, and it can run off some common data analytics file formats.

fsiefken

10 days ago

DuckDB is an open-source column-oriented Relational Database Management System (RDBMS). It's designed to provide high performance on complex queries against large databases in embedded configuration.

It has transparent compression built-in and has support for natural language queries. https://buckenhofer.com/2025/11/agentic-ai-with-duckdb-and-s...

"DICT FSST (Dictionary FSST) represents a hybrid compression technique that combines the benefits of Dictionary Encoding with the string-level compression capabilities of FSST. This approach was implemented and integrated into DuckDB as part of ongoing efforts to optimize string storage and processing performance." https://homepages.cwi.nl/~boncz/msc/2025-YanLannaAlexandre.p...

linhns

10 days ago

1 reply

Not the author here. I’m not sure about DuckDB, but SQLite allows you to simply use a file as a database and for archiving, it’s really helpful. One file, that’s it.

cobolcomesback

10 days ago

1 reply

DuckDB does as well. A super simplified explanation of duckdb is that it’s sqlite but columnar, and so is better for analytics of large datasets.

formerly_proven

10 days ago

1 reply

The schema is this: items(id INTEGER PRIMARY KEY, type TEXT, time INTEGER, by TEXT, title TEXT, text TEXT, url TEXT

Doesn't scream columnar database to me.

embedding-shape

10 days ago

1 reply

At a glance, that is missing (at least) a `parent` or `parent_id` attribute which items in HN can have (and you kind of need if you want to render comments), see http://hn.algolia.com/api/v1/items/46436741

agolliver

10 days ago

[delayed]

3eb7988a1663

10 days ago

1 reply

While I suspect DuckDB would compress better, given the ubiquity of SQLite, it seems a fine standard choice.

peheje

10 days ago

1 reply

the data is dominated by big unique TEXT columns, unsure how that can much compress better when grouped - but would be interesting to know

3eb7988a1663

9 days ago

I was thinking more the numeric columns which have pre-built compression mechanisms to handle incrementing columns or long runs of identical values. For sure less total data than the text, but my prior is that the two should perform equivalently on the text, so the better compression on numbers should let duckdb pull ahead.

I had to run a test for myself, and using sqlite2duckdb (no research, first search hit), and using randomly picked shard 1636, the sqlite.gz was 4.9MB, but the duckdb.gz was 3.7MB.

The uncompressed sizes favor sqlite, which does not make sense to me, so not sure if duckdb keeps around more statistics information. Uncompressed sqlite 12.9MB, duckdb 15.5MB

jacquesm

10 days ago

1 reply

Maybe it got nuked by MS? The rest of their repo's are up.

keepamovin

10 days ago

1 reply

Hey jacquesm! No, I just forgot to make it public.

BUT I did try to push the entire 10GB of shards to GitHub (no LFS, no thanks, money), and after the 20 minutes compressing objects etc, "remote hang up unexpectedly"

To be expected I guess. I did not think GH Pages would be able to do this. So have been repeating:

  wrangler pages deploy docs --project-name static-news --commit-dirty=true

on changes and first time CF Pages user here, much impressed!

jacquesm

10 days ago

1 reply

Pretty neat project. I never thought you could do this in the first place, very much inspiring. I've made a little project that stores all of its data locally but still runs in the browser to protect against take downs and because I don't think you should store your precious data online more than you have to, eventually it all rots away. Your project takes this to the next level.

keepamovin

10 days ago

1 reply

Thanks, bud, that means a lot! Would like to see your versions of the data stored offline idea, it's very cool.

jacquesm

9 days ago

1 reply

pianojacq.com

It's super simple, really, far less impressive than what you've built there.

keepamovin

8 days ago

That's really cool, man. The music notation is beautiful. I hit play but couldn't get it to progress past the first note. Maybe I need to plug in a midi keyboard? It would be cool if I could "play" with my ASCII keyboard.

Listen was nice. That's really cool, actually. I encourage you to do it.

keepamovin

10 days ago

i forgot to set repo to public. Fixed now

wslh

10 days ago

1 reply

Is this updated regularly? 404 on GitHub as the other comment.

scsh

10 days ago

The BQ dataset is only ~17GB and the free tier of BQ lets you query 1TB per month. If you're not doing select * on every query you should be able to do a lot with that.

abixb

10 days ago

2 replies

Wonder if you could turn this into a .zim file for offline browsing with an offline browser like Kiwix, etc. [0]

I've been taking frequent "offline-only-day" breaks to consolidate whatever I've been learning, and Kiwix has been a great tool for reference (offline Wikipedia, StackOverflow and whatnot).

[0] https://kiwix.org/en/the-new-kiwix-library-is-available/

keepamovin

10 days ago

Can you submit a PR to repo with a script that does that? Sounds cool.

Barbing

10 days ago

Oh this should TOTALLY be available to those who are scrolling through sources on the Kiwix app!

fao_

10 days ago

5 replies

> Community, All the HN belong to you. This is an archive of hacker news that fits in your browser.

> 20 years of HN arguments and beauty, can be yours forever. So they'll never die. Ever. It's the unkillable static archive of HN and it's your hands

I'm really sorry to have to ask this, but this really feels like you had an LLM write it?

rantingdemon

10 days ago

2 replies

Why do you say that?

sundarurfriend

10 days ago

2 replies

[delayed]

JavGull

10 days ago

1 reply

With the em dashes I see you. But at this point idrc so long as it reads well. Everyone uses spell check…

naikrovek

10 days ago

3 replies

I add em dashes to everything I write now, solely to throw people who look for them off. Lots of editors add them automatically when you have two sequential dashes between words — a common occurrence, like that one. And this is is Chrome on iOS doing it automatically.

Ooh, I used “sequential”, ooh, I used an em dash. ZOMG AI IS COMING FOR US ALL

fao_

9 days ago

I also use Em-Dashes, this is about how weird the thing is tonally

Barbing

10 days ago

Ya—in fact, globally replaced on iOS (sent from Safari)

Also for reference: “this shortcut can be toggled using the switch labeled 'Smart Punctuation' in General > Keyboard settings.”

3eb7988a1663

10 days ago

Anyone demonstrating above a high-school vocabulary/reading level is obviously a machine.

deadbabe

10 days ago

2 replies

Sometimes I want to write more creatively, but then worry I’ll be accused of being an LLM. So I dumb it down. Remove the colorful language. Conform.

ssl-3

10 days ago

1 reply

Fuck 'em.

Always write what you want, however you want to write it. If some reader somewhere decides to be judgemental because of — you know — an em dash or an X/Y comparison or a complement or some other thing that they think pins you down as a bot, then that's entirely their own problem. Not yours.

They observe the reality that they deserve.

deadbabe

10 days ago

You’re absolutely right. It’s not my problem, it’s their problem.

catlifeonmars

10 days ago

[delayed]

fao_

9 days ago

It feels like a LLM doing it's usual "gushing out appreciations"

naikrovek

10 days ago

1 reply

> I'm really sorry to have to ask this, but this really feels like you had an LLM write it?

Ending with a question mark doesn’t make your sentence a question. You didn’t ask anything. You stated an opinion and followed it with a question mark.

If you intended to ask a question about the text being written by AI, no, you don’t have to ask that.

I am so damn tired of the “that didn’t happen” and the “AI did that” people when there is zero evidence of either being true.

These people are the most exhausting people I have ever encountered in my entire life.

jacquesm

10 days ago

You're right. Unfortunately they are also more and more often right.

jesprenj

10 days ago

1 reply

I doubt it. "hacker news" spelled lowercase? comma after "beauty"? missing "in" after "it's"? i doubt an LLM would make such syntax mistakes. it's just good writing, that's also possible these days.

fao_

9 days ago

> it's just good writing, that's also possible these days.

As someone reskilling into being a writer, I really do not think that is "good writing".

Insanity

10 days ago

Even if so, would it have mattered? The point is showing off the SQLite DB.

But it didn’t read LLM generated IMO.

walthamstow

10 days ago

There's a thing in soccer at the moment where a tackle looks fine in realtime but when the video referee shows it to the onpitch referee, they show the impact in slo-mo over and over again and it always looks worse.

I wonder if there's something like this going on here. I never thought it was LLM on first read, and I still don't, but when you take snippets and point at them it makes me think maybe they are

tevon

10 days ago

1 reply

The link seems to be down, was it taken down?

scsh

10 days ago

Probably just forgot to make it public.

yupyupyups

10 days ago

1 hour passed and it's already nuked?

Thank you btw

asdefghyk

10 days ago

2 replies

How much space is needed? ...for the data .... Im wondering if it would work on a tablet? ....

asdefghyk

10 days ago

FYI I did NOT see the size info in the title. Impossible to edit / delete my comment now ........

keepamovin

10 days ago

~9GB gzipped.

zX41ZdbW

10 days ago

1 reply

The query tab looks quite complex with all these content shards: https://hackerbook.dosaygo.com/?view=query

I have a much simpler database: https://play.clickhouse.com/play?user=play#U0VMRUNUIHRpbWUsI...

embedding-shape

10 days ago

1 reply

Does your database also runs offline/locally in the browser? Seems to be the reason for the large number of shards.

zX41ZdbW

10 days ago

You can run it locally, but it is a client-server architecture, which means that something has to run behind the browser.

Paul-E

10 days ago

2 replies

That's pretty neat!

I did something similar. I build a tool[1] to import the Project Arctic Shift dumps[2] of reddit into sqlite. It was mostly an exercise to experiment with Rust and SQLite (HN's two favorite topics). If you don't build a FTS5 index and import without WAL, import of every reddit comment and submission takes a bit over 24 hours and produces a ~10TB DB.

SQLite offers a lot of cool json features that would let you store the raw json and operate on that, but I eschewed them in favor of parsing only once at load time. THat also lets me normalize the data a bit.

I find that building the DB is pretty "fast", but queries run much faster if I immediately vacuum the DB after building it. The vacuum operation is actually slower than the original import, taking a few days to finish.

[1] https://github.com/Paul-E/Pushshift-Importer

[2] https://github.com/ArthurHeitmann/arctic_shift/blob/master/d...

s_ting765

10 days ago

1 reply

[delayed]

Paul-E

10 days ago

I haven't tested that, so I'm not sure if it would work. The import only inserts rows, it doesn't delete, so I don't think that is the cause of fragmentation. I suspect this line in the vacuum docs:

> The VACUUM command may change the ROWIDs of entries in any tables that do not have an explicit INTEGER PRIMARY KEY.

means SQLite does something to organize by rowid and that this is doing most of the work.

Reddit post/comment IDs are 1:1 with integers, though expressed in a different base that is more friendly to URLs. I map decoded post/comment IDs to INTEGER PRIMARY KEYs on their respective tables. I suspect the vacuum operation sorts the tables by their reddit post ID and something about this sorting improves tables scans, which in turn helps building indices quickly after standing up the DB.

Xyra

10 days ago

Holy cow, I didn't know getting reddit was that straightforward. I am building public readonly-SQL+vector databases optimized for exploring high-quality public commons with Claude Code (https://exopriors.com/scry), I so cannot wait until some funding source comes in and I can upgrade to a $1500/month Hetzner server and pay the ~$1k to embed all that.

simonw

10 days ago

9 replies

Don't miss how this works. It's not a server-side application - this code runs entirely in your browser using SQLite compiled to WASM, but rather than fetching a full 22GB database it instead uses a clever hack that retrieves just "shards" of the SQLite database needed for the page you are viewing.

I watched it in the browser network panel and saw it fetch:

  https://hackerbook.dosaygo.com/static-shards/shard_1636.sqlite.gz
  https://hackerbook.dosaygo.com/static-shards/shard_1635.sqlite.gz
  https://hackerbook.dosaygo.com/static-shards/shard_1634.sqlite.gz

As I paginated to previous days.

It's reminiscent of that brilliant SQLite.js VFS trick from a few years ago: https://github.com/phiresky/sql.js-httpvfs - only that one used HTTP range headers, this one uses sharded files instead.

The interactive SQL query interface at https://hackerbook.dosaygo.com/?view=query asks you to select which shards to run the query against, there are 1636 total.

nextaccountic

10 days ago

9 replies

Is there anything more production grade built around the same idea of HTTP range requests like that sqlite thing? This has so much potential

simonw

10 days ago

2 replies

There was a UK government GitHub repo that did something interesting with this kind of trick against S3 but I checked just now and the repo is a 404. Here are my notes about what it did: https://simonwillison.net/2025/Feb/7/sqlite-s3vfs/

simonw

10 days ago

3 replies

I recovered it from https://archive.softwareheritage.org/browse/origin/directory... and pushed a fresh copy to GitHub here:

https://github.com/simonw/sqlite-s3vfs

This comment was helpful in figuring out how to get a full Git clone out of the heritage archive: https://news.ycombinator.com/item?id=37516523#37517378

AceJohnny2

10 days ago

1 reply

didn't you do something similar for Datasette, Simon?

simonw

10 days ago

1 reply

Nothing smart with HTTP range requests yet - I have https://lite.datasette.io which runs the full Python server app in the browser via WebAssembly and Pyodide but it still works by fetching the entire SQLite file at once.

AceJohnny2

10 days ago

oh! I must've been confused with your TIL where you explained this technique

https://simonwillison.net/2021/May/2/hosting-sqlite-database...

QuantumNomad_

10 days ago

1 reply

I also have a locally cloned copy of that repo from when it was on GitHub. Same latest commit as your copy of it.

From what I see in GitHub in your copy of the repo, it looks like you don’t have the tags.

Do you have the tags locally?

If you don’t have the tags, I can push a copy of the repo to GitHub too and you can get the tags from my copy.

simonw

10 days ago

1 reply

I don't have the tags! It would be awesome if you could push that.

QuantumNomad_

10 days ago

1 reply

Uploaded here:

https://github.com/Quantum-Nomad/sqlite-s3vfs

simonw

10 days ago

1 reply

Thanks for that, though actually it turns out I had them after all - I needed to run:

  git push --tags origin

QuantumNomad_

10 days ago

All the better :)

bspammer

10 days ago

1 reply

[delayed]

socialcommenter

9 days ago

From reading the TIL, it doesn't appear as if Simon used LLM for a large portion of what he did; only the initial suggestion to check the archive, and the web tool to make his process reproducible. Also, if you read the script from his chat with Claude code, the prompt really does the heavy lifting.

Sure, the LLM fills in all the boilerplate and makes an easy-to-use, reproducible tool with loads of documentation, and credit for that. But is it not more accurate to say that Simon is absurdly efficient, LLM or sans LLM? :)

billywhizz

10 days ago

i played around with this a while back. you can see a demo here. it also lets you pull new WAL segments in and apply them to the current database. never got much time to go any further with it than this.

https://just.billywhizz.io/sqlite/demo/#https://raw.githubus...

Humphrey

10 days ago

2 replies

Yes — PMTiles is exactly that: a production-ready, single-file, static container for vector tiles built around HTTP range requests.

I’ve used it in production to self-host Australia-only maps on S3. We generated a single ~900 MB PMTiles file from OpenStreetMap (Australia only, up to Z14) and uploaded it to S3. Clients then fetch just the required byte ranges for each vector tile via HTTP range requests.

It’s fast, scales well, and bandwidth costs are negligible because clients only download the exact data they need.

https://docs.protomaps.com/pmtiles/

simonw

10 days ago

2 replies

PMTiles is absurdly great software.

Humphrey

10 days ago

1 reply

I know right! I'd never heard of HTTP Range requests until PMTiles - but gee it's an elegant solution.

keepamovin

10 days ago

Hadn't seen PMTiles before, but that matches the mental model exactly! I chose physical file sharding over Range Requests on a single db because it felt safer for 'dumb' static hosts like CF. - less risk of a single 22GB file getting stuck or cached weirdly. Maybe it would work

hyperbolablabla

10 days ago

2 replies

My only gripe is that the tile metadata is stored as JSON, which I get is for compatibility reasons with existing software, but for e.g. a simple C program to implement the full spec you need to ship a JSON parser on top of the PMTiles parser itself.

seg_lol

10 days ago

1 reply

A JSON parser is less than a thousand lines of code.

Diti

9 days ago

2 replies

And where most of CPU time will be wasted in, if you care about profiling/improving responsiveness.

monerozcash

9 days ago

At that point you're just io bound, no? I can easily parse json at 100+GB/s on commodity hardware, but I'm gonna have a much harder time actually delivering that much data to parse.

keepamovin

9 days ago

What's a better way?

keepamovin

9 days ago

How would you store it?

nextaccountic

10 days ago

1 reply

That's neat, but.. is it just for cartographic data?

I want something like a db with indexes

jtbaker

10 days ago

1 reply

Look into using duckdb with remote http/s3 parquet files. The parquet files are organized as columnar vectors, grouped into chunks of rows. Each row group stores metadata about the set it contains that can be used to prune out data that doesn’t need to be scanned by the query engine. https://duckdb.org/docs/stable/guides/performance/indexing

LanceDB has a similar mechanism for operating on remote vector embeddings/text search.

It’s a fun time to be a dev in this space!

nextaccountic

8 days ago

1 reply

> Look into using duckdb with remote http/s3 parquet files. The parquet files are organized as columnar vectors, grouped into chunks of rows. Each row group stores metadata about the set it contains that can be used to prune out data that doesn’t need to be scanned by the query engine. https://duckdb.org/docs/stable/guides/performance/indexing

But, when using this on frontend, are portions of files fetched specifically with http range requests? I tried to search for it but couldn't find details

jtbaker

5d ago

Yes, you should be able to see the byte range requests and 206 responses from an s3 compatible bucket or http server that supports those access patterns.

6510

10 days ago

1 reply

I want to see a bittorrent version :P

nextaccountic

10 days ago

Maybe webtorrent-based?

mootothemax

10 days ago

This is pretty much well what is so remarkable about parquet files; not only do you get seekable data, you can fetch only the columns you want too.

I believe that there are also indexing opportunities (not necessarily via eg hive partitioning) but frankly - am kinda out of my depth pn it.

omneity

10 days ago

I tried to implement something similar to optimize sampling semi-random documents from (very) large datasets on Huggingface, unfortunately their API doesn't support range requests well.

ericd

10 days ago

[delayed]

tlarkworthy

10 days ago

Parquet/iceberg

taywrobel

10 days ago

Here’s a very relevant blog post about doing so in a static webpage using GitHub pages, including an interactive demo - https://phiresky.github.io/blog/2021/hosting-sqlite-database...

Disclaimer: I work at GitHub but am unaffiliated with that blog.

__turbobrew__

10 days ago

gdal vsis3 dynamically fetches chunks of rasters from s3 using range requests. It is the underlying technology for several mapping systems.

There is also a file format to optimize this https://cogeo.org/

ncruces

10 days ago

2 replies

A read-only VFS doing this can be really simple, with the right API…

This is my VFS: https://github.com/ncruces/go-sqlite3/blob/main/vfs/readervf...

And using it with range requests: https://pkg.go.dev/github.com/ncruces/go-sqlite3/vfs/readerv...

And having it work with a Zstandard compressed SQLite database, is one library away: https://pkg.go.dev/github.com/SaveTheRbtz/zstd-seekable-form...

pdyc

10 days ago

1 reply

this does not caches the data right? it would always fetch from network? by any chance do you know of solution/extension that caches the data it would make it so much more efficient.

ncruces

10 days ago

The package I'm using in the HTTP example can be configured to cache data: https://github.com/psanford/httpreadat?tab=readme-ov-file#ca...

But, also, SQLite caches data; you can simply increase the page cache.

keepamovin

10 days ago

1 reply

Your page is served over sqlitevfs with Range queries? Let's try this here.

ncruces

9 days ago

1 reply

I did a similar VFS in Go. It doesn't run client-side in a browser.

But you can use it (e.g.) in a small VPS to access a multi-TB database directly from S3.

keepamovin

9 days ago

That is cool. Maybe i look at the go code

meander_water

10 days ago

1 reply

I love this so much, on my phone this is much faster than actual HN (I know it's only a read-only version).

Where did you get the 22GB figure from? On the site it says:

> 46,399,072 items, 1,637 shards, 8.5GB, spanning Oct 9, 2006 to Dec 28, 2025

amitmahbubani

10 days ago

2 replies

> Where did you get the 22GB figure from?

The HN post title (:

keepamovin

10 days ago

22GB is non-gzipped.

meander_water

10 days ago

Hah, well that's embarrassing

sodafountan

10 days ago

1 reply

The GitHub page is no longer available, which is a shame because I'm really interested in how this works.

How was the entirety of HN stored in a single SQLite database? In other words, how was the data acquired? And how does the page load instantly if there's 22GB of data having to be downloaded to the browser?

keepamovin

10 days ago

1 reply

You can see it now, forgot to make it public.

- 1. download_hn.sh - bash script that queries BigQuery and saves the data to *.json.gz

- 2. etl-hn.js - does the sharding and ID -> shard map, plus the user stats shards.

- 3. Then either npx serve docs or upload to CloudFlare Pages.

The ./toool/s/predeploy-checks.sh script basically runs the entire pipeline. You can do it unattended with AUTO_RUN=true

sodafountan

10 days ago

Awesome, I'll take a look

maxloh

10 days ago

1 reply

I am curios why they don't use a single file and HTTP Range Requests instead. PMTiles (a distribution of OpenStreetMap) uses that.

keepamovin

10 days ago

This would be a neat idea to try. Want to add a PR? Bench different "hackends" to see how DuckDB, SQLite shards, or range queries perform?

keepamovin

10 days ago

1 reply

Thanks! I'm glad you enjoyed the sausage being made. There's a little easter egg if you click on the compact disc icon.

oblosys

9 days ago

1 reply

I almost got tricked into trying to figure out what was Easter eggy about August 9 2015 :-) There's a clarifying tooltip on the link, but it is mostly obscured by the image's "Archive" title attribute.

keepamovin

9 days ago

1 reply

Oh, shit that was the problem! You solved the bug! I was trying to figure out why the right tooltip didn't display. A linked wrapped in an image wrapped in an easter egg! Or something. Ha, thank you. Will fix :)

oblosys

9 days ago

Happy to help!

keepamovin

9 days ago

A recent change is I added date spans to the shard checboxes on query view so it's easier to zero dates you want if you have that in mind. Because if your copy isn't local all those network pulls take a while.

The sequence of shards you saw when you paginated to days is faciliated by the static-manifest which maps HN item ID ranges to shards, and since IDs are increasing and a pretty good proxy of time (a "HN clock"), we can also map the shards that we divided by ID, to time spans their items cover. An in memory table sorted by time is created from the manifest on load so we can easily look up which shard we need when you pick a day.

Funnily enough, this system was thrown off early on by a handful of "ID/timestamp" outliers in the data: items with weird future timestamps (offset by a couple years), or null timestamps. To cleanse our pure data from this noise, and restore proper adjacent-in-time shard cuts we just did a 1/99 percentile grouping and discarded the outliers leaving shards with sensible 'effective' time spans.

Sometimes we end up fetching two shards when you enter a new day because some items' comments exist "cross shard". We needed another index for that and it lives in cross-shard-index.bin which is just a list of 4-byte item IDs that have children in more than 1 shard (2-bytes), which occurs when people have the self-indulgence to respond to comments a few days after a post has died down ;)

Thankfully HN imposes a 2 week horizon for replies so there aren't that many cross-shard comments (those living outside the 2-3 days span of most, recent, shards). But I think there's still around 1M or so, IIRC.

dzhiurgis

9 days ago

Is it possible to implement search this way?

tehlike

10 days ago

Vfs support is amazing.

sieep

10 days ago

4 replies

What a reminder on how text is so much more efficient than video, its crazy! Could you imagine the same amount of knowledge (or dribble) but in video form? I wonder how large that would be.

ivanjermakov

10 days ago

1 reply

Average high quality 1080p60 video has bitrate of 5Mbps, which is equivalent to 120k English words per second. With average English speech being 150wpm, we end up with text being 50 thousand times more space efficient.

sieep

9 days ago

Cheers to you for doing the math, I hope 2026 is excellent for you!

fsiefken

10 days ago

2 replies

one could use a video llm to generate the video, diagrams or the stills automatically based on the text. except when it's boardgames playthroughs or programming i just transcribe to text, summarise and read youtube video's.

Barbing

10 days ago

1 reply

Can be nice to pull a raw transcript and have it formatted as HTML (formatting/punctuation fixes applied).

Best locally of course to avoid “I burned a lake for this?” guilt.

fsiefken

10 days ago

yes, yt-dlp can download the transcript, and if it's not available i can get the audio file and run it through parakeet locally.

deskamess

10 days ago

1 reply

How do you read youtube videos? Very curious as I have been wanting to watch PDF's scroll by slowly on a large TV. I am interested in the workflow of getting a pdf/document into a scrolling video format. These days NotebookLM may be an option but I am curious if there is something custom. If I can get it into video form (mp4) then I can even deliver it via plex.

fsiefken

10 days ago

I use yt-dlp to download the transcript, and if it's not available i can get the audio file and run it through parakeet locally. Then I have the plain text, which could be read out loud (kind of defeating the purpose), but perhaps at triple speed with a computer voice that's still understandble at that speed. I could also summarize it with an llm. With pandoc or typst I can convert to single column or mult column pdf to print or watch on tv or my smart glasses. If I strip the vowels and make the font smaller I can fit more!

jacquesm

10 days ago

1 reply

That's what's so sad about youtube. 20 minute videos to encode a hundred words of usable content to get you to click on a link. The inefficiency is just staggering.

Rendello

10 days ago

1 reply

Youtube can be excellent for explanations. A picture's worth a thousand words, and you can fit a lot of decent pictures in a 20 minute video. The signal-to-noise can be high, of course.

ComputerGuru

9 days ago

1 reply

Unfortunately even the videos that do contain helpful imagery are still dominated by huge sections of low entropy.

For example, one of the most useful applications of video over text is appliance or automotive repair, but the ideal format would be an article interspersed with short video sections, not a video with a talking head and some ~static shaky cam taking up most of the time as the individual drones on about mostly unrelated topics or unimportant details yet you can’t skip past it in case there is something actually pertinent covered in that time.

Rendello

9 days ago

Ay, there's the rub. Professional video makes tend to be pushed into making videos for a more general audience, and niche topics are left to first-timers who haven't developed video-making skills and (tend to) go on and on.

I've produced a few videos, and I was shocked at how difficult it was to be clear. I have the same problem with writing, but at least it's restricted in a way video making isn't. There's so many ways to make a video about something, and most of them are wrong!

keepamovin

10 days ago

Right? 20 years, probably 10s millions of human hours of interactions, and it’s only as much as a couple DVDs.

sirjaz

10 days ago

1 reply

This would be awesome as a cross platform app.

keepamovin

10 days ago

Good idea. HN.exe

zkmon

10 days ago

2 replies

Similar to Single-page applications (SPA), single-table application (STA) might become a thing. Just a shard a table on multiple keys and serve the shards as static files, provided that the data is Ok to share, similar to sharing static html content.

jesprenj

10 days ago

1 reply

do you mean single database? it'd be quite hard if not impossible to make applications using a single table (no relations). reddit did it though, they have a huge table of "things" iirc.

mburns

10 days ago

1 reply

[delayed]

rplnt

10 days ago

And the important lesson from that the k/v-like aspect of it. That the "schema" is horizontal (is that a thing?) and not column-based. But I actually only read it on their blog IIRC and never even got the full details - that there's still a third ID column. Thanks for the link.

jhd3

10 days ago

[The Baked Data architectural pattern](https://simonwillison.net/2021/Jul/28/baked-data/)

yread

10 days ago

6 replies

I wonder how much smaller it could get with some compression. You could probably encode "This website hijacks the scrollbar and I don't like it" comments into just a few bits.

Rendello

10 days ago

1 reply

The hard-coded dictionary wouldn't be much stranger than Brotli's:

https://news.ycombinator.com/item?id=27159506

maxbond

10 days ago

You can use a BPE variant like SentencePiece to identify these patterns rather than hard coding them.

rossant

10 days ago

Guilty.

keepamovin

8 days ago

Dear it is already compressed using G zip – nine for every SQLlight shard and manifest

22 GB is uncompressed and compressed the entire things about 9 GB

hamburglar

10 days ago

It might be a neat experiment to use ai to produce canonicalized paraphrasings of HN arguments so they could be compared directly and compress well.

keepamovin

9 days ago

It is gzipped (sqilte, bin, json manifests) from 22GB s raw to ~9GB.

jacquesm

10 days ago

That's at least 45%, then you can leave out all of my comments and you're left with only 5!

spit2wind

10 days ago

1 reply

This is pretty neat! The calendar didn't work well for me. I could only seem to navigate by month. And when I selected the earliest day (after much tapping), nothing seemed to be updated.

Nonetheless, random access history is cool.

keepamovin

10 days ago

Cna you let me know? I'm sure there's some weirdness lurking there and I want to smooth it out. Calendar is essential.

Sn0wCoder

10 days ago

2 replies

Site does not load on Firefox console error says 'Uncaught (in promise) TypeError: can't access property "wasm", sqlite3 is null'

Guess its common knowledge that SharedArrayBuffer (SQLite wasm) does not work with FF due to Cross-Origin Attacks (i just found out ;).

Once the initial chunk of data loads the rest load almost instantly on Chrome. Can you please fix the GitHub link (current 404) would like to peak at the code. Thank you!

keepamovin

10 days ago

1 reply

Damn. Will try to fix for FF

Sn0wCoder

9 days ago

1 reply

Strange now the first few days load (getting a new error) 'Ignoring inability to install OPFS sqlite3_vfs: Cannot install OPFS: Missing SharedArrayBuffer and/or Atomics. The server must emit the COOP/COEP response headers to enable those. See https://sqlite.org/wasm/doc/trunk/persistence.md#coop-coep'

But when go back to the 26th none of the shards will load, error out.

Using Windows 11, FF 146.0.1

Since you tested it seems its just a me problem and thanks for fixing the GitHub link

keepamovin

9 days ago

No I've seen that error too, on Safari. I think it's related to the wasm being sent with wrong headers. CF pages _headers file should be ensuring correctness. Can you try busting your cache (or wait for a new Dec 29 Data dump version coming in a couple minutes), or from incognito to see if that fixes the issue? It's possible an earlier version had stale headers or sth. Idk.

coder543

10 days ago

[delayed]

layer8

10 days ago

1 reply

Apparently the comment counts are only the top-level comments?

It would be nice for the thread pages to show a comment count.

keepamovin

10 days ago

Yes, because comments in a thread can span shards. It’s just a bit too heavy to add comment counts of an entire thread. So I give a low bound ha ha

dspillett

10 days ago

1 reply

Is there a public dump of the data anywhere that this is based upon, or have they scraped it themselves?

Such as DB might be entertaining to play with, and the threadedness of comments would be useful for beginners to practise efficient recursive queries (more so than the StackExchange dumps, for instance).

thomasmarton

10 days ago

While not a dump per se, there is an API where you can get HN data programmatically, no scraping needed.

https://github.com/HackerNews/API

dmarwicke

10 days ago

22gb for mostly text? tried loading the site, it's pretty slow. curious how the query performance is with this much data in sqlite

57 more comments available on Hacker News

Resources