Similar series / Readers also like
Updated by user · 29 days ago · 9 min read
Near the bottom of every series page there are two recommendation rails: **Similar series** and **Readers also like**. They look alike, but they answer different questions and are built from completely different data. Here's how each one works. Both are still in **beta**. ## TL;DR **Similar series** turns each series into a "fingerprint" of its tags, where rarer, more specific tags count for more, then finds the series whose fingerprints point the same way. **Readers also like** ignores tags entirely and instead looks at MangaBaka libraries: if a lot of readers keep both this series and another, the two get linked. Both lists are then cleaned up (no five editions of the same story), reordered for variety, and adjusted by the thumbs-up / thumbs-down feedback people leave on recommendations. ## Two different questions | Rail | The question it answers | Built from | | --------------------- | ------------------------------------------- | -------------------------- | | **Similar series** | "What else is _like_ this one?" | The series' own tags | | **Readers also like** | "What do people who read this _also_ read?" | MangaBaka reader libraries | The first is about the _content_ of a series. The second is about _reading habits_, and it can connect two series that share almost no tags, as long as the same people tend to keep both. ## How "Similar series" works Series-to-series similarity is built on tags. Three ideas do most of the work. ### 1. Every series becomes a weighted tag vector We take all of a series' tags and turn them into a long list of numbers, a _vector_, that captures what the series is "about". Three things decide how much each tag counts: - **How rare the tag is.** A tag that sits on a handful of series tells you far more than one that sits on a hundred thousand (like _Romance_). Rare tags get more weight, common ones get less. This is a standard technique called _IDF_ (inverse document frequency). - **How specific the tag is.** Our tags form a tree, from broad categories down to narrow sub-tropes. The deeper (more specific) the tag, the more it counts. Broad genres are treated as a flat label: they help, but they don't dominate. - **How central the tag is to _that_ series.** Each tag on a series carries a relevance level, from a _core_ theme down to an _incidental_ background detail. A core tag counts for much more than an incidental one, so the same tag can shape one series strongly and barely register on another. The upshot is that each series is described mostly by its _distinctive_, _central_ tags, not the generic ones nearly every title shares. The tags don't have to sit directly on a series, either. Tags live in a tree, so a series also counts toward the _broader_ categories above each of its tags: tag something with a narrow sub-trope and its parent themes count too. That's why two series can match on a shared theme even when only one carries the exact narrow tag. (Organizational tags — the housekeeping labels that aren't about content — don't count.) ### 2. We find the nearest neighbours With every series reduced to a vector, "similar" becomes "points in the same direction". We measure the angle between two series' vectors (_cosine similarity_) and pull the closest matches. This runs against a purpose-built index, so we can search hundreds of thousands of series in a few milliseconds instead of comparing against every title one by one. ### 3. We adjust the raw matches Pure tag overlap isn't the whole story, so a few more signals adjust the ranking: - **Same author or artist** gets a boost (and a _Same author_ badge). - **Same format** (manga, manhwa, novel, and so on) counts for a little, so a manga leans slightly toward other manga. - **Editorial "related series"**, hand-curated relationships, get the biggest boost (and a _Related_ badge). - **Shared-tag count** matters on top of the angle: two series with _lots_ of tags in common edge ahead of two that merely happen to line up. - **Recognizability.** Because the fingerprint favours rare tags, it can surface very obscure exact-matches. A small popularity weighting keeps well-known, on-theme series from getting buried under niche ones. - **Community feedback.** A thumbs-up/down on a recommendation feeds back into that pairing's ranking, and a series that gets broadly thumbs-downed across the site is pushed down everywhere. - **Lower-quality candidates** (fan works / doujinshi, cancelled titles) are pushed down a little rather than hidden. ### When a series barely has any tags The vector trick needs a reasonable number of tags to work. For a thinly-tagged series (fewer than ~12 tags) the math degrades and tends to return other thinly-tagged junk. So below that threshold we switch approach entirely: we take the source's most _defining_ tags and look for series where those same tags are central, ranked by recognizability. A series with no tags at all can't be matched this way (see the caveats). ## Keeping the list useful Before you see it, the list gets two cleanups: - **Edition collapse.** The exact same story you're already looking at shouldn't show up as a "similar" result just because it's a different edition, language, or the novel version, so those are dropped. And a single franchise is capped at **2** entries, so an eight-volume spin-off web doesn't flood the rail. - **Variety.** Without a diversity pass you'd often get the same five near-identical titles in a row. We re-order to keep strong matches near the top while breaking up clusters, so the list stays varied without drifting off-theme. Each result also explains itself: hover the info icon and you'll see how many **tags in common** there are (the top 10 are listed), plus _Same author_ or _Related_ where they apply. The label — _Very similar_, _Similar_, _Somewhat similar_ — is just a bucket on the shared-tag count (10+, 5+, or fewer). Spoiler tags are the exception: they still count toward the match, but we keep them out of the explanation so a recommendation never gives away a plot twist. Every Similar series rail also has a **Start a mix** link that blends this series (and any others you add) into a custom recommendation feed. ## How "Readers also like" works This rail ignores tags completely. It's **collaborative filtering**: "people who read this also read that". We look across every MangaBaka reader's library and count how often two series show up _together_. A series someone **completed** or **rated highly** counts as a strong positive; one they **dropped** counts against the pairing. The more readers who keep both, and the more positively they keep them, the stronger the link. Two guardrails keep the pairings meaningful: - A pairing needs at least **3 different readers** who keep both series before we'll show it. That's a cold-start gate: brand-new or rarely-shelved series won't have a "readers also like" list yet. - One very active reader can't dominate. We only count each reader's top entries, so a 5,000-title library doesn't outweigh everyone else. As with the other rail, the results then get the same community-feedback adjustment and variety pass. ## How fresh is it - **Tag fingerprints** are rebuilt **daily**, right after we recompute how rare each tag is. New tags, merged tags, and newly-tagged series flow in within a day. - **Reader-library pairings** are rebuilt **weekly**, with a light top-up **every 15 minutes** that folds in recent library changes so the popular pairings don't go stale between full rebuilds. So if you tag a series today, its **Similar series** reflects that tomorrow; if a bunch of people shelve it this week, **Readers also like** catches up over the following days. ## Caveats & things to keep in mind > [!IMPORTANT] > Both rails are still in **beta**, and "similar" means _thematically related_, not _good_. A series can be a near-perfect tag match and still not be to your taste. These are starting points, not verdicts. A few more things to be upfront about: - **It's only as good as the tags.** Similar series rests entirely on tag quality. A sparsely- or poorly-tagged series gets weak or odd matches, and a completely untagged one gets _no_ Similar series at all. If a result looks wrong, better tags are the fix, and they show up the next day. - **New series start cold.** A title nobody has tagged or shelved yet won't have much on either rail. This fills in as the catalogue and community grow. - **Our reader signal is still young.** _Readers also like_ leans on the size of the MangaBaka community, which is small today, so for most series it's thinner than _Similar series_. It gets better as more people build libraries. - **Your filters apply.** Your content-rating settings and any tags you've blocked are respected, so two people can see slightly different lists for the same series. - **It moves.** Both rails are recomputed on a schedule and shift with new tags, new readers, and feedback, even on days nothing about the series itself changed. ## FAQ **What's the difference between the two rails?** _Similar series_ compares the series themselves, by their tags. _Readers also like_ compares _people_, by what they keep in their libraries. One can suggest things the other never would. **Why is some unrelated-looking series showing up as "similar"?** Almost always tags. The two share more (or rarer) tags than it looks, or one of them is tagged in a way that doesn't match how you'd describe it. Hover the info icon to see the shared tags, and fix the tags if they're off. **Why don't I see any Similar series at all?** The series probably has too few tags, or none. Tag it, and the rail populates within a day. **Why is "Readers also like" empty when "Similar series" isn't?** It needs at least three readers who keep this series _and_ another one in common. Newer or niche series often haven't crossed that bar yet. **Can I influence the recommendations?** Yes. Tag series accurately (that drives _Similar series_), and use the thumbs-up/down on each recommendation — your feedback re-ranks that pairing and feeds the global signal. **Are these lists personalized to me?** Only lightly. The matches are a property of the _series_, so everyone sees the same Similar series and Readers also like for a given title. They aren't tailored to your taste the way a personal feed would be. The only personal part is that your content-rating preferences and any blocked tags are applied on top, so two people with different settings can see slightly different lists. Your own library doesn't change them. **How often does it update?** Similar series: daily. Readers also like: weekly, with a 15-minute incremental top-up.