A library data story · goodbooks-10k

What people kept almost choosing

A library knows what was checked out — never what was picked up, considered, and quietly put back. This is an attempt to recover that missing shelf: 912,705 real "want to read" shelvings, checked against whether each reader ever came back to finish the book.

10,000books catalogued
53,424readers
5.98Mcompleted checkouts
912,705almost-chose shelvings

01 — The shelf

The books people reached for most, and finished least

For every book someone added to their "want to read" shelf, we checked whether that same reader ever logged a rating for it. If not, it goes on this shelf — not abandoned, exactly, but not finished either.

still on the shelf Cover of The Book Thief

The Book Thief

Markus Zusak

2,762 almost chose 0.36% finished
still on the shelf Cover of All the Light We Cannot See

All the Light We Cannot See

Anthony Doerr

1,964 almost chose 0.15% finished
still on the shelf Cover of Catch-22

Catch-22

Joseph Heller

1,839 almost chose 0.05% finished
still on the shelf Cover of 1984

1984

George Orwell, Erich Fromm, Celâl Üster

1,804 almost chose 0.44% finished
still on the shelf Cover of The Kite Runner

The Kite Runner

Khaled Hosseini

1,765 almost chose 0.11% finished
still on the shelf Cover of Life of Pi

Life of Pi

Yann Martel

1,710 almost chose 0.41% finished
still on the shelf Cover of Miss Peregrine’s Home for Peculiar Children (Miss Peregrine’s Peculiar Children, #1)

Miss Peregrine’s Home for Peculiar Children (Miss Peregrine’s Peculiar Children, #1)

Ransom Riggs

1,647 almost chose 0.18% finished
still on the shelf Cover of Slaughterhouse-Five

Slaughterhouse-Five

Kurt Vonnegut Jr.

1,608 almost chose 0.00% finished

02 — The pattern

Some genres get finished. Others just get intended.

Grouping every "almost chose" by genre surfaces a clean split: contemplative, single-sitting nonfiction is shelved with intention and rarely returned to. Fast, plot-driven fiction converts that same intention into a finished book far more often.

Mostly still on the shelf

Memoir0.02%
Philosophy0.02%
Biography0.03%
Childrens0.04%
Science0.08%
History0.10%

Usually followed through

Contemporary0.47%
Romance0.38%
Mystery0.38%
Fantasy0.36%
Young Adult0.34%
Urban Fantasy0.30%

03 — The model

A recommender that knows the difference

Item-item latent factors (truncated SVD over 5.98M ratings) find books that read alike. Blending in each candidate's follow-through rate nudges the ranking toward books people who picked them up actually finished — a small hedge against recommending another "almost chose."

If you almost chose

Catch-22

Joseph Heller

→ ranked by similarity, weighted toward books readers actually finish

Cover of Slaughterhouse-Five

Slaughterhouse-Five

0.00% follow-through

Cover of One Flew Over the Cuckoo's Nest

One Flew Over the Cuckoo's Nest

0.00% follow-through

Cover of A Clockwork Orange

A Clockwork Orange

0.39% follow-through

Cover of Brave New World

Brave New World

0.00% follow-through

04 — Method & limits

What this signal is, and isn't

The proxy. "Almost chose" = a reader shelved the book on Goodreads' "want to read" list but never logged a rating for it in this dataset. It's a real, directly observed behavioral signal — not a simulation — but it's a proxy, not a fact: some of these readers may finish the book tomorrow.

Why the rates look small. The underlying rating file is a capped sample of each reader's history, so absolute follow-through percentages run low across the board. What's trustworthy is the relative comparison between books and genres under identical sampling, not the raw percentage.

Genres are crowd-sourced. Labels come from the highest-count Goodreads folksonomic tag per book, filtered to a genre-like vocabulary — not a controlled taxonomy.

Goodreads isn't a library. Shelving and rating behavior is a reasonable public stand-in for a hold-and-checkout system, not an exact one. The dataset is goodbooks-10k, built from real Goodreads activity.