01 — The shelf
The books people reached for most, and finished least
For every book someone added to their "want to read" shelf, we checked whether that same reader ever logged a rating for it. If not, it goes on this shelf — not abandoned, exactly, but not finished either.
The Book Thief
All the Light We Cannot See
Catch-22
1984
The Kite Runner
Life of Pi
Miss Peregrine’s Home for Peculiar Children (Miss Peregrine’s Peculiar Children, #1)
Slaughterhouse-Five
02 — The pattern
Some genres get finished. Others just get intended.
Grouping every "almost chose" by genre surfaces a clean split: contemplative, single-sitting nonfiction is shelved with intention and rarely returned to. Fast, plot-driven fiction converts that same intention into a finished book far more often.
Mostly still on the shelf
Usually followed through
03 — The model
A recommender that knows the difference
Item-item latent factors (truncated SVD over 5.98M ratings) find books that read alike. Blending in each candidate's follow-through rate nudges the ranking toward books people who picked them up actually finished — a small hedge against recommending another "almost chose."
If you almost chose
Catch-22
Joseph Heller
→ ranked by similarity, weighted toward books readers actually finish
Slaughterhouse-Five
0.00% follow-through
One Flew Over the Cuckoo's Nest
0.00% follow-through
A Clockwork Orange
0.39% follow-through
Brave New World
0.00% follow-through
04 — Method & limits
What this signal is, and isn't
The proxy. "Almost chose" = a reader shelved the book on Goodreads' "want to read" list but never logged a rating for it in this dataset. It's a real, directly observed behavioral signal — not a simulation — but it's a proxy, not a fact: some of these readers may finish the book tomorrow.
Why the rates look small. The underlying rating file is a capped sample of each reader's history, so absolute follow-through percentages run low across the board. What's trustworthy is the relative comparison between books and genres under identical sampling, not the raw percentage.
Genres are crowd-sourced. Labels come from the highest-count Goodreads folksonomic tag per book, filtered to a genre-like vocabulary — not a controlled taxonomy.
Goodreads isn't a library. Shelving and rating behavior is a reasonable public stand-in for a hold-and-checkout system, not an exact one. The dataset is goodbooks-10k, built from real Goodreads activity.