SiteLiner — internal site hygiene, and the duplicate content penalty that was never real
Crawl your own site for internal duplicate content, broken links and thin pages. What the free scan actually covers, where it stops, and when to reach for a heavier crawler.
- Duplicate Content
- Broken Links
- Site Audit
- Internal Linking
- Technical SEO
- Publisher
- Indigo Stream Technologies
- Type
- Site Audit Crawler
- Pricing
- Freemium
- Reviewed
- 19 September 2026
Quick verdict
Use when
- You have just migrated, redesigned or moved a site and want to know what broke before search engines find it
- You suspect the same content exists on several of your own URLs and want to see which pages compete with each other
- You want to check a site without creating an account, installing anything, or handing over an email address
- You need an answer in minutes and the site is small enough that a partial crawl still covers the important pages
- You maintain a site whose content is server-rendered, so a crawler that reads plain HTML sees everything that matters
- You want a cheap first pass before deciding whether a heavier paid crawler is worth the setup
Skip when
- Your site is built as a client-rendered application, where a crawler without JavaScript will see very little of the real content
- You need to know whether your text appears elsewhere on the web, which is a different job from finding duplication inside your own site
- The site is large and you need every page crawled rather than the most prominent ones
- You want continuous monitoring with alerts, so that a problem is caught the week it appears rather than the month you remember to look
- You need performance metrics, Core Web Vitals or rendering diagnostics, which are outside what this tool reports
- You need to crawl behind a login or script a crawl as part of a pipeline
SiteLiner vs Screaming Frog vs Ahrefs vs Search Console
Tap a dimension to focus
Pricing
SiteLiner leads- SiteLinerThis page
- A free tier with no account required, capped by how many pages are crawled and how often a single site can be re-checked
- The free allowance covers the most prominent pages chosen by internal link structure rather than the whole site
- A paid tier raises the ceiling to a large page count, removes the frequency limit, allows crawl controls and keeps previous reports
- Signing up for the paid tier is itself free; you buy prepaid scan credits as you need them rather than committing to a subscription
- The per-page metering makes a one-off cleanup cheap and continuous monitoring comparatively expensive, which is the opposite of most tools here
- A free tier capped by how many URLs it will crawl, which is generous for small sites and limiting for real ones
- A paid licence removes the cap and unlocks scheduled crawls, integrations and the deeper configuration options
- Licensed per user rather than per site, so an agency pays per seat rather than per client
- No free tier for the audit itself; it is part of a paid subscription to the wider platform
- Tiers scale by how much of the platform you use — credits, projects and seats — rather than by crawl volume alone
- The audit is effectively included once you are paying for the suite, which makes the marginal cost zero if you already subscribe
- Free, with no paid tier at all
- Coverage limits are set by the platform, not by you, and there is nothing to buy to raise them
- For most site owners this is the first and only tool they need for indexation questions
Preview of SiteLiner - not the live app. Confirm details on the official site.
Does this overview help you decide?
—
(—)
Learn more
Details below the decision summary—features, workflow, and scope notes.
What is SiteLiner?
What it costs
- Free tier
- Yes
- Pricing summary
- The pricing model here is metered by pages crawled rather than by month, which is unusual in this category and shapes when the tool is worth using. The free tier needs no account at all: you type a domain and it crawls a capped number of your most prominent pages, with the limit being how often the same site can be checked. That cap is the part people misread. The pages are chosen by internal link structure, not by importance or recency, so the ones it picks are the ones that are easiest to reach from your navigation. That is a reasonable proxy for a small site and a poor one for a large one, and it means the free scan can miss exactly the pages most likely to be causing trouble — the deeply nested ones and the near-orphans. It also will not see anything rendered by JavaScript, so on a client-rendered application the report can look almost empty and be misleading rather than merely partial. Above the free tier, signing up costs nothing and you buy scan credits in advance, sized to the number of pages you want crawled. There is no recurring commitment, which makes a one-off cleanup after a migration genuinely cheap, and it makes continuous monitoring a poor fit — there is no scheduler and no alerting, so the cost model and the use case point the same way. Note also that this sits alongside the vendor’s other service rather than replacing it: that one checks a page against the wider web and is priced separately, and if you arrived looking for plagiarism detection rather than internal duplication, that is the product you actually want. Since the crawl allowance and credit pricing both change, confirm the current figures on the official product page before you plan around them.
Reviewed on 19 September 2026 · SiteLiner — free and premium products
What SiteLiner includes
Internal duplicate content, which is the reason it exists
It reports where the same or near-identical content appears across your own URLs, which is the question the sibling plagiarism service could not answer. Internal duplication is usually accidental rather than deliberate: a platform migration that left the old pages live, product descriptions reused across variants, a page reworked without retiring the previous version. Seeing the pairs listed is what turns a vague suspicion into a decision about which URL should be canonical.
Broken links, internal and outbound
It reports links that lead nowhere, including outbound links to pages that have since disappeared. Outbound rot is the part people neglect, because nothing in analytics tells you that a page you cited three years ago is gone. It is also a user-facing problem rather than only a search one, since a dead link is a dead end in the middle of a reading experience.
An internal link map
The crawl records how your pages link to one another, which shows up both as a report and implicitly in how it chooses which pages to scan. That structure is worth seeing directly: pages with almost no incoming internal links are hard for both crawlers and readers to reach, and a page you care about being found that nobody links to is a common and easily fixed oversight.
Page weight, word counts and shared content
Each page is reported with its size, its word count, and how much of its content is shared with other pages on the site. The last of these is the more interesting one, because it separates genuine duplication from boilerplate — navigation, footers and legal text are identical everywhere by necessity, and a report that flagged those as problems would be noise. Being able to see the distinction is what makes the duplicate-content number actionable.
A report you do not have to configure
There is no crawl setup, no rules and no dashboard to learn: the output is a set of tables on a page, and you read it. That is a deliberate design choice and it is the tool’s main advantage over a professional crawler, which can do far more but requires you to decide what it should do first. The audience here is someone who wants an answer, not someone who wants a tool to configure.
Crawl controls above the free tier
The paid tier lets you direct which parts of the site are scanned and raises the page ceiling substantially, which matters on a large site where the default "most prominent pages" selection would otherwise decide your audit for you. It also keeps previous reports, so a run can be compared with an earlier one rather than being lost when the tab closes.
No account needed to start
The free scan runs without signup, which sounds trivial and is the feature that most affects how the tool gets used. Removing the account step removes the reason to postpone, and it also means nothing about your scan is tied to a user record. For a quick sanity check after a deploy, that friction difference is the whole value proposition.
How to use an internal duplication report without chasing the wrong problem
Start from what you just changed, not from a blank audit
The most productive time to run this is immediately after a migration, a redesign, a platform move or a bulk content edit, because you have a specific hypothesis about what might have broken and the report can confirm or clear it in minutes. Running it with no trigger tends to produce a list of long-standing issues that were never causing measurable harm, and a long list with no priority is not a finding.
Check whether content is server-rendered before you trust the numbers
This is the single most important check, because the failure is silent: a crawler that does not execute JavaScript will see a client-rendered application as mostly empty. That does not produce an error, it produces a report showing very little content, which reads as "your site is fine" when the truth is "your site was invisible to this crawler". If your pages need JavaScript to show their content, use a crawler that renders, and treat this tool as unable to answer the question.
Separate intentional duplication from accidental
Not all repetition is a problem. Alternates built for tracking, print-friendly versions and language variants are deliberate, and flagging them as defects trains you to ignore the report. What you are hunting for is duplication you did not choose: an old page still live after a rewrite, the same description across many product variants, or a template change that published two copies. Decide which URLs should be canonical and consolidate, rather than deleting content that is doing a job.
Ignore the boilerplate, act on the content
Navigation, footers, cookie notices and legal text are identical across every page by definition, and a report that surfaces them is telling you nothing. The signal is duplication in the body content — the part that is supposed to be unique to each page. Reading a shared-content percentage without that distinction produces busywork such as rewriting a footer to make a number go down.
Fix the internal link structure while you are looking at it
The same crawl tells you how your pages connect, and the pages with no internal links pointing at them are both harder for search engines to prioritise and easier for readers to miss. Adding a few contextual links from relevant pages is usually a smaller change than any content rewrite and often a larger effect, and it is the corrective action available from this report that does not require deleting anything.
Know when to stop using the free scan and reach for a real crawler
A partial crawl of the most prominent pages is the right first question and the wrong last one. If the site is large, if the issue you are chasing lives in a deep section, if you need JavaScript rendering, or if you need to re-run the audit on a schedule and be told when something regresses, you have outgrown a domain field and a few minutes. The honest summary of this tool is that it tells you whether to take the problem seriously — and the answer sometimes is that you need a heavier instrument.
Who SiteLiner is for
Anyone who has just migrated or rebuilt a site
The clearest fit. Migrations are where duplication and dead links are created in bulk — old URLs left live, redirects pointing at the wrong targets, content republished under new paths without retiring the originals — and there is a short window where catching it is cheap. A tool with no install and no account is the one you will actually run in that window rather than the one you mean to set up later.
Developers and maintainers doing a periodic sanity pass
A structural check of your own site is developer work, not marketing work: broken links, orphaned pages with no internal links, page weight, and accidental duplicate routes are all things that come out of code and templates. Running it as part of a release checklist is realistic precisely because it costs nothing and takes minutes, which is not true of a configured crawl.
Small sites whose owner is also the maintainer
On a site of a few dozen pages the free allowance covers essentially everything that matters, and the plain-HTML crawler sees all of it. The absence of a dashboard is an advantage here rather than a limitation, because there is no wrong way to read a table of pages and their issues, and no subscription to justify for a check you run twice a year.
People evaluating whether a duplicate content problem exists at all
Before commissioning work or buying a tool, it is worth knowing whether the thing you are worried about is real. This answers that in one step, which makes it a good way to avoid solving an imagined problem — and equally good at confirming a real one, since a list of pages with near-identical bodies settles the question quickly.
Teams who need a first pass before a paid tool is justified
The heavier crawlers and the audit suites both cost money and setup, and neither is worth it if a five-minute scan shows nothing structural. Using this as the triage step keeps the paid instrument for the sites and moments that need it, which for a small team is the difference between auditing properly and not auditing at all.
When SiteLiner is the right pick
Product and platform notes
- A plain HTML crawler, and that decides which sites it suits
- It fetches pages and reads the markup without executing JavaScript. On a conventionally rendered site that is all you need and it is fast. On a client-rendered application it sees a shell, and the report will understate the site rather than warn you about it — a silent failure rather than a visible one. Check this before trusting any finding on a modern framework site.
- It audits one site against itself, not against the web
- This is the distinction that confuses the most people, because the tool comes from the makers of the best-known plagiarism service and the names invite the assumption that they do the same thing. They do not: this one finds content repeated across your own URLs, and checking whether your text has been copied elsewhere is the separate service. If plagiarism is what you need, this is the wrong tool and the footnote below points to the right one.
- The free crawl is the most prominent pages, not all of them
- Selection is by internal link structure, so the pages it examines are the easiest to reach from your navigation, and the cap applies to how often the same site can be re-scanned. On a small site that is effectively a full audit. On a large one it is a sample, and a sample weighted toward well-linked pages — which means the deep and hard-to-reach pages, often the source of both duplication and orphan problems, are the least likely to appear in it.
- There is no duplicate content penalty, and the page should say so
- Search engines have stated for years that they do not penalise sites for duplicate content. What duplication actually costs is crawl efficiency, a forced choice between candidate URLs, and diluted signals — real problems with real fixes, but not a punishment, and the distinction changes what you should do about it. Deliberate alternates used for tracking or printing are fine; accidental second copies of the same page need consolidating. A page that implied a penalty to make an audit sound urgent would be repeating the folklore rather than reporting the position.
- No monitoring, no alerting, no scheduling
- It reports on the moment you asked. There is no recurring crawl and nothing that emails you when something regresses, which is the boundary between it and the paid audit suites. That is the right design for a check you run deliberately and the wrong one for a property you need watched, and it is worth knowing which of those you actually have before choosing.
- What it does not do
- It does not render JavaScript, so it cannot see client-rendered content. It does not check your text against the wider web. It does not crawl behind a login or accept scripting as part of a pipeline. It does not report performance, Core Web Vitals or rendering diagnostics. It does not monitor or alert, and it does not cover every page on a large site on the free tier. It also does not tell you which duplicate to keep — that remains a judgement about your content, not a number a crawler can produce.