SiteLiner — internal site hygiene, and the duplicate content penalty that was never real

Crawl your own site for internal duplicate content, broken links and thin pages. What the free scan actually covers, where it stops, and when to reach for a heavier crawler.

  • Duplicate Content
  • Broken Links
  • Site Audit
  • Internal Linking
  • Technical SEO
Publisher
Indigo Stream Technologies
Type
Site Audit Crawler
Pricing
Freemium
Reviewed
19 September 2026
Official site

Quick verdict

Use when

  • You have just migrated, redesigned or moved a site and want to know what broke before search engines find it
  • You suspect the same content exists on several of your own URLs and want to see which pages compete with each other
  • You want to check a site without creating an account, installing anything, or handing over an email address
  • You need an answer in minutes and the site is small enough that a partial crawl still covers the important pages
  • You maintain a site whose content is server-rendered, so a crawler that reads plain HTML sees everything that matters
  • You want a cheap first pass before deciding whether a heavier paid crawler is worth the setup

Skip when

  • Your site is built as a client-rendered application, where a crawler without JavaScript will see very little of the real content
  • You need to know whether your text appears elsewhere on the web, which is a different job from finding duplication inside your own site
  • The site is large and you need every page crawled rather than the most prominent ones
  • You want continuous monitoring with alerts, so that a problem is caught the week it appears rather than the month you remember to look
  • You need performance metrics, Core Web Vitals or rendering diagnostics, which are outside what this tool reports
  • You need to crawl behind a login or script a crawl as part of a pipeline

Try instead

If the problem is how the site performs, how it is tracked, or what to do about the content itself, those live here too.

SiteLiner vs Screaming Frog vs Ahrefs vs Search Console

The first thing to get straight is that these four are not all doing the same job, and two of them are not competitors at all. This product is the fastest route to an answer: you type a domain and get a report on internal duplication, broken links and page weight without signing up or installing anything, and the report is generated by crawling plain HTML. Screaming Frog is the professional desktop crawler and the one most technical users eventually graduate to — it runs locally on your machine rather than in someone else's cloud, gives you a level of configuration the browser tools do not attempt, and handles JavaScript rendering so client-rendered sites are actually visible to it. The trade is that it is an application to install and learn, and its free tier is capped by URL count. Ahrefs is a cloud crawler that lives inside a broad paid SEO subscription: it schedules re-crawls, keeps a running health score, tracks issues over time and adds performance data, which is what you want if site health is something you monitor rather than something you check. Google Search Console is the odd one out and belongs in the comparison anyway, because it is free, official and tells you something none of the crawlers can — how Google actually sees your site, including which pages it indexed, which it excluded and which URLs it hit and failed to find. What it will not do is crawl your site for internal duplication. So the sensible workflow is layered rather than exclusive: use this product or Search Console for a quick read, move to the desktop crawler when you need every page and precise control, and pay for the suite when you want the audit to run on a schedule and shout at you. Choose on whether website health is an event or an ongoing responsibility, because that is the question that actually separates these tools.

Tap a dimension to focus

Pricing

SiteLiner leads
  • SiteLinerThis page
    • A free tier with no account required, capped by how many pages are crawled and how often a single site can be re-checked
    • The free allowance covers the most prominent pages chosen by internal link structure rather than the whole site
    • A paid tier raises the ceiling to a large page count, removes the frequency limit, allows crawl controls and keeps previous reports
    • Signing up for the paid tier is itself free; you buy prepaid scan credits as you need them rather than committing to a subscription
    • The per-page metering makes a one-off cleanup cheap and continuous monitoring comparatively expensive, which is the opposite of most tools here
    • A free tier capped by how many URLs it will crawl, which is generous for small sites and limiting for real ones
    • A paid licence removes the cap and unlocks scheduled crawls, integrations and the deeper configuration options
    • Licensed per user rather than per site, so an agency pays per seat rather than per client
    • No free tier for the audit itself; it is part of a paid subscription to the wider platform
    • Tiers scale by how much of the platform you use — credits, projects and seats — rather than by crawl volume alone
    • The audit is effectively included once you are paying for the suite, which makes the marginal cost zero if you already subscribe
    • Free, with no paid tier at all
    • Coverage limits are set by the platform, not by you, and there is nothing to buy to raise them
    • For most site owners this is the first and only tool they need for indexation questions
  1. SiteLiner — official site
  2. SiteLiner — free and premium products
  3. SiteLiner — frequently asked questions
  4. Copyscape — the related web-wide service
  5. Screaming Frog SEO Spider — official site
  6. Ahrefs — site audit
  7. Google Search Console — official page

Preview of SiteLiner - not the live app. Confirm details on the official site.

Does this overview help you decide?

(—)

Learn more

Details below the decision summary—features, workflow, and scope notes.

What is SiteLiner?

Crawl your own site for internal duplicate content, broken links and thin pages. What the free scan actually covers, where it stops, and when to reach for a heavier crawler.

What it costs

A free scan with no account required and a monthly limit per site, plus a paid tier that removes the frequency limit and is bought as prepaid scan credits rather than a subscription.
Free tier
Yes
Pricing summary
The pricing model here is metered by pages crawled rather than by month, which is unusual in this category and shapes when the tool is worth using. The free tier needs no account at all: you type a domain and it crawls a capped number of your most prominent pages, with the limit being how often the same site can be checked. That cap is the part people misread. The pages are chosen by internal link structure, not by importance or recency, so the ones it picks are the ones that are easiest to reach from your navigation. That is a reasonable proxy for a small site and a poor one for a large one, and it means the free scan can miss exactly the pages most likely to be causing trouble — the deeply nested ones and the near-orphans. It also will not see anything rendered by JavaScript, so on a client-rendered application the report can look almost empty and be misleading rather than merely partial. Above the free tier, signing up costs nothing and you buy scan credits in advance, sized to the number of pages you want crawled. There is no recurring commitment, which makes a one-off cleanup after a migration genuinely cheap, and it makes continuous monitoring a poor fit — there is no scheduler and no alerting, so the cost model and the use case point the same way. Note also that this sits alongside the vendor’s other service rather than replacing it: that one checks a page against the wider web and is priced separately, and if you arrived looking for plagiarism detection rather than internal duplication, that is the product you actually want. Since the crawl allowance and credit pricing both change, confirm the current figures on the official product page before you plan around them.

Reviewed on 19 September 2026 · SiteLiner — free and premium products

What SiteLiner includes

The capabilities as the product documents them, read against what each one is for on a real site.
  • Internal duplicate content, which is the reason it exists

    It reports where the same or near-identical content appears across your own URLs, which is the question the sibling plagiarism service could not answer. Internal duplication is usually accidental rather than deliberate: a platform migration that left the old pages live, product descriptions reused across variants, a page reworked without retiring the previous version. Seeing the pairs listed is what turns a vague suspicion into a decision about which URL should be canonical.

  • Broken links, internal and outbound

    It reports links that lead nowhere, including outbound links to pages that have since disappeared. Outbound rot is the part people neglect, because nothing in analytics tells you that a page you cited three years ago is gone. It is also a user-facing problem rather than only a search one, since a dead link is a dead end in the middle of a reading experience.

  • An internal link map

    The crawl records how your pages link to one another, which shows up both as a report and implicitly in how it chooses which pages to scan. That structure is worth seeing directly: pages with almost no incoming internal links are hard for both crawlers and readers to reach, and a page you care about being found that nobody links to is a common and easily fixed oversight.

  • Page weight, word counts and shared content

    Each page is reported with its size, its word count, and how much of its content is shared with other pages on the site. The last of these is the more interesting one, because it separates genuine duplication from boilerplate — navigation, footers and legal text are identical everywhere by necessity, and a report that flagged those as problems would be noise. Being able to see the distinction is what makes the duplicate-content number actionable.

  • A report you do not have to configure

    There is no crawl setup, no rules and no dashboard to learn: the output is a set of tables on a page, and you read it. That is a deliberate design choice and it is the tool’s main advantage over a professional crawler, which can do far more but requires you to decide what it should do first. The audience here is someone who wants an answer, not someone who wants a tool to configure.

  • Crawl controls above the free tier

    The paid tier lets you direct which parts of the site are scanned and raises the page ceiling substantially, which matters on a large site where the default "most prominent pages" selection would otherwise decide your audit for you. It also keeps previous reports, so a run can be compared with an earlier one rather than being lost when the tab closes.

  • No account needed to start

    The free scan runs without signup, which sounds trivial and is the feature that most affects how the tool gets used. Removing the account step removes the reason to postpone, and it also means nothing about your scan is tied to a user record. For a quick sanity check after a deploy, that friction difference is the whole value proposition.

How to use an internal duplication report without chasing the wrong problem

The loop that works, and the five places this kind of audit goes wrong.
  1. Start from what you just changed, not from a blank audit

    The most productive time to run this is immediately after a migration, a redesign, a platform move or a bulk content edit, because you have a specific hypothesis about what might have broken and the report can confirm or clear it in minutes. Running it with no trigger tends to produce a list of long-standing issues that were never causing measurable harm, and a long list with no priority is not a finding.

  2. Check whether content is server-rendered before you trust the numbers

    This is the single most important check, because the failure is silent: a crawler that does not execute JavaScript will see a client-rendered application as mostly empty. That does not produce an error, it produces a report showing very little content, which reads as "your site is fine" when the truth is "your site was invisible to this crawler". If your pages need JavaScript to show their content, use a crawler that renders, and treat this tool as unable to answer the question.

  3. Separate intentional duplication from accidental

    Not all repetition is a problem. Alternates built for tracking, print-friendly versions and language variants are deliberate, and flagging them as defects trains you to ignore the report. What you are hunting for is duplication you did not choose: an old page still live after a rewrite, the same description across many product variants, or a template change that published two copies. Decide which URLs should be canonical and consolidate, rather than deleting content that is doing a job.

  4. Ignore the boilerplate, act on the content

    Navigation, footers, cookie notices and legal text are identical across every page by definition, and a report that surfaces them is telling you nothing. The signal is duplication in the body content — the part that is supposed to be unique to each page. Reading a shared-content percentage without that distinction produces busywork such as rewriting a footer to make a number go down.

  5. Fix the internal link structure while you are looking at it

    The same crawl tells you how your pages connect, and the pages with no internal links pointing at them are both harder for search engines to prioritise and easier for readers to miss. Adding a few contextual links from relevant pages is usually a smaller change than any content rewrite and often a larger effect, and it is the corrective action available from this report that does not require deleting anything.

  6. Know when to stop using the free scan and reach for a real crawler

    A partial crawl of the most prominent pages is the right first question and the wrong last one. If the site is large, if the issue you are chasing lives in a deep section, if you need JavaScript rendering, or if you need to re-run the audit on a schedule and be told when something regresses, you have outgrown a domain field and a few minutes. The honest summary of this tool is that it tells you whether to take the problem seriously — and the answer sometimes is that you need a heavier instrument.

Who SiteLiner is for

The site owners and situations this tool actually maps onto.
  • Anyone who has just migrated or rebuilt a site

    The clearest fit. Migrations are where duplication and dead links are created in bulk — old URLs left live, redirects pointing at the wrong targets, content republished under new paths without retiring the originals — and there is a short window where catching it is cheap. A tool with no install and no account is the one you will actually run in that window rather than the one you mean to set up later.

  • Developers and maintainers doing a periodic sanity pass

    A structural check of your own site is developer work, not marketing work: broken links, orphaned pages with no internal links, page weight, and accidental duplicate routes are all things that come out of code and templates. Running it as part of a release checklist is realistic precisely because it costs nothing and takes minutes, which is not true of a configured crawl.

  • Small sites whose owner is also the maintainer

    On a site of a few dozen pages the free allowance covers essentially everything that matters, and the plain-HTML crawler sees all of it. The absence of a dashboard is an advantage here rather than a limitation, because there is no wrong way to read a table of pages and their issues, and no subscription to justify for a check you run twice a year.

  • People evaluating whether a duplicate content problem exists at all

    Before commissioning work or buying a tool, it is worth knowing whether the thing you are worried about is real. This answers that in one step, which makes it a good way to avoid solving an imagined problem — and equally good at confirming a real one, since a list of pages with near-identical bodies settles the question quickly.

  • Teams who need a first pass before a paid tool is justified

    The heavier crawlers and the audit suites both cost money and setup, and neither is worth it if a five-minute scan shows nothing structural. Using this as the triage step keeps the paid instrument for the sites and moments that need it, which for a small team is the difference between auditing properly and not auditing at all.

When SiteLiner is the right pick

Duplicate content is one of the few technical SEO conversations where the folklore is louder than the facts, so it is worth being precise about what a tool like this is actually for. Search engines do not penalise a site for having the same text on two URLs. What duplication does is waste crawl budget, force the search engine to choose between candidate URLs, and split whatever signals would otherwise concentrate on one page — which is a real problem with a real cost, just not a punishment. That reframing matters because it tells you when to care. If you have deliberately created alternates for tracking or for print, the duplication is intentional and harmless. If the same product description appears across twelve variants, or a reworked page left the old one live, or a platform change introduced a second copy of everything, then you have a canonicalisation problem and this is the class of tool that finds it. Where this particular product earns its place is the entry price and the entry effort. There is no account, no install, no configuration, and no waiting for a report to be queued: it is a domain field and a few minutes, which means it gets used at the moment you think of it rather than being saved for a project. Its limits are honest and worth respecting. The crawl is partial by design, it does not render JavaScript, it will not see behind a login, and it reports on the moment you asked rather than watching over time. Add that it compares pages within your site and not against the web, and you have a tool with a clear edge and a clear boundary — which is exactly what makes it a good first pass before you decide whether the heavier crawler or the paid suite is warranted. Use it when the question is "what did I just break" or "is my own site competing with itself", and move on to something with a scheduler when the answer needs to be monitored rather than discovered.

Product and platform notes

What it is, what it will not do, and the facts worth verifying at the source.
A plain HTML crawler, and that decides which sites it suits
It fetches pages and reads the markup without executing JavaScript. On a conventionally rendered site that is all you need and it is fast. On a client-rendered application it sees a shell, and the report will understate the site rather than warn you about it — a silent failure rather than a visible one. Check this before trusting any finding on a modern framework site.
It audits one site against itself, not against the web
This is the distinction that confuses the most people, because the tool comes from the makers of the best-known plagiarism service and the names invite the assumption that they do the same thing. They do not: this one finds content repeated across your own URLs, and checking whether your text has been copied elsewhere is the separate service. If plagiarism is what you need, this is the wrong tool and the footnote below points to the right one.
The free crawl is the most prominent pages, not all of them
Selection is by internal link structure, so the pages it examines are the easiest to reach from your navigation, and the cap applies to how often the same site can be re-scanned. On a small site that is effectively a full audit. On a large one it is a sample, and a sample weighted toward well-linked pages — which means the deep and hard-to-reach pages, often the source of both duplication and orphan problems, are the least likely to appear in it.
There is no duplicate content penalty, and the page should say so
Search engines have stated for years that they do not penalise sites for duplicate content. What duplication actually costs is crawl efficiency, a forced choice between candidate URLs, and diluted signals — real problems with real fixes, but not a punishment, and the distinction changes what you should do about it. Deliberate alternates used for tracking or printing are fine; accidental second copies of the same page need consolidating. A page that implied a penalty to make an audit sound urgent would be repeating the folklore rather than reporting the position.
No monitoring, no alerting, no scheduling
It reports on the moment you asked. There is no recurring crawl and nothing that emails you when something regresses, which is the boundary between it and the paid audit suites. That is the right design for a check you run deliberately and the wrong one for a property you need watched, and it is worth knowing which of those you actually have before choosing.
What it does not do
It does not render JavaScript, so it cannot see client-rendered content. It does not check your text against the wider web. It does not crawl behind a login or accept scripting as part of a pipeline. It does not report performance, Core Web Vitals or rendering diagnostics. It does not monitor or alert, and it does not cover every page on a large site on the free tier. It also does not tell you which duplicate to keep — that remains a judgement about your content, not a number a crawler can produce.

Frequently Asked Questions

Quick answers about this tool—open a question to read more.