Back to guides

Content Health: Auditing a Library Instead of a Page

Content health audits the library rather than the page: depth, cadence, internal links, authorship, schema and date integrity across everything you publish.

Most technical checks look at one page. Content health looks at the library: how much you have published, how recently, how deep it goes, how well the pieces connect to each other, and whether the metadata around them is consistent enough to be believed. It needs its own audit because library-level problems are invisible page by page. Every article can be individually decent while the set holds 3 pieces competing for the same query, 40 posts nothing links to, dates a plugin bumped without the text changing, or a publishing cadence that stopped 14 months ago. None of that shows up when you read a single URL. Decay is the pattern worth watching. Content that ranked well 2 years ago drifts down as the topic moves and the article stays where it was. The signals that predict decay are measurable well before the traffic drops: thin evidence, stale references, missing authorship, no recent internal links pointing in. The checks below read your published set as a portfolio, covering depth, uniqueness, heading structure, internal link topology, topical clustering, authorship consistency, schema completeness and date integrity. Read the results as one diagnosis rather than as a list of separate defects.

No content program detected

Why it matters

A visible, fresh, and differentiated content program helps the site compete beyond its homepage and core product pages.

How we check it

The crawl sample was checked for article roots, Article or BlogPosting schema, article Open Graph type, and indexable article-like pages.

How to fix it

Create and expose a crawlable article, guide, resource, or case-study library that targets the questions your buyers research before they contact you.

Article discoverability is weak

How we check it

Article pages were checked for sitemap presence and internal-link discoverability.

How to fix it

Include article URLs in the XML sitemap and link important articles from crawlable non-article pages.

Editorial content share is measured, not scored

How we check it

The sampled crawl and sitemap URLs were measured for article coverage, and the editorial share is reported as a descriptive percentage.

How to fix it

Treat this share as context rather than a target: page volume carries no threshold here and does not affect the score. Let the scored content findings — discovery, freshness, uniqueness, and article structure — decide where the library needs work.

Content program is stale

How we check it

Article publish and update dates were extracted from JSON-LD, time elements, and page metadata.

How to fix it

Publish or substantially update at least one article every quarter, and expose datePublished/dateModified signals on the page.

Editorial publishing cadence is measured, not scored

How we check it

Unique editorial resources published in the last 365 days were counted. Translations of the same resource count once; no thin-content quality claim is implied.

How to fix it

Treat this count as context rather than a target: publishing frequency carries no threshold here and does not affect the score. Let the freshness and content-decay findings, which are scored, set how often you publish.

Editorial article length is measured, not scored

How we check it

Visible article text was tokenized and measured after boilerplate cleanup, and the median length is reported as a descriptive count.

How to fix it

Treat the median length as context rather than a target: no word-count threshold applies here and length does not affect the score. Judge each article on whether it answers its reader's question with original explanation, first-hand evidence, examples, and practical next steps.

Near-duplicate articles detected

How we check it

Article pages were compared with exact shingle Jaccard similarity after boilerplate cleanup.

How to fix it

Consolidate near-duplicate articles, use canonical URLs where consolidation is not possible, and make each retained page target a distinct search intent.

Article heading structure is weak

How we check it

Article pages were checked for H2 and H3 structure.

How to fix it

Use clear H2 sections to organize each article around the questions and subtopics searchers need answered.

Topic and intent clusters are mapped, not scored

Why it matters

Observed clusters help editors see portfolio shape, overlap, and linking opportunities without treating inferred demand as fact.

How we check it

Sampled content pages were grouped by observed topic terms and intent markers without inferring search demand.

How to fix it

Treat this map as context rather than a verdict: clusters are counted, overlap and connectivity between them are not evaluated, and nothing here affects the score. Use the observed clusters as an editorial review map — validate audience questions, consolidate material overlap, and connect genuinely related pages.

Content portfolio classified

How we check it

Crawled content was separated into editorial articles, programmatic resource collections, other article-like pages, and general pages.

How to fix it

Keep editorial articles and programmatic resource collections on separate URL patterns so each type can be judged on its own merits and a large template collection cannot stand in for editorial depth.

Programmatic templates need stronger differentiation

Why it matters

Repeated templates need page-specific substance so each URL contributes distinct, verifiable value.

How we check it

Repeated page families were compared for boilerplate, entity-specific value, and within-family similarity.

How to fix it

Add entity-specific facts, first-party evidence, distinct answers, and useful local or item-level details to repeated templates; consolidate pages that cannot justify a distinct purpose.

Source and original-evidence inventory is measured, not scored

Why it matters

Source, methodology, media, and data inventories show where readers can verify or extend the claims on a page.

How we check it

Sampled content pages were inventoried for external sources, methodology sections, first-party media, data assets, tables, and figures.

How to fix it

Treat this inventory as context rather than a verdict: sources, methodology pages, media, and data assets are counted, not judged, and the count does not affect the score. Where editorial claims depend on outside facts or original work, expose the relevant source links, methodology, media, tables, figures, or downloadable data.

Editorial authorship needs strengthening

Why it matters

Clear author identity and profile links help readers and search systems understand who created editorial content and why that source is credible.

How we check it

Editorial Article markup was checked for an author name and a stable profile URL or @id.

How to fix it

Identify each editorial author in Article markup and link to a stable profile page or @id that explains the author's relevant experience.

Author identity signals conflict

Why it matters

Consistent bylines and stable author identities make editorial responsibility easier to verify.

How we check it

Visible bylines were compared with Article author names, stable profile URLs, and sameAs identities.

How to fix it

Align visible bylines with Article author markup and use one stable profile URL or @id for each person across the site.

Article schema coverage is thin

How we check it

Detected article pages were checked for Article or BlogPosting structured data.

How to fix it

Add complete Article or BlogPosting JSON-LD to article pages, including headline, author, datePublished, dateModified, and image where available.

Article schema fields are incomplete

Why it matters

Complete Article markup connects the headline, author, image, and publication date to the visible page without requiring search systems to infer those relationships.

How we check it

Structured content pages were checked for headline, author, image, and datePublished fields.

How to fix it

Complete Article or BlogPosting markup with an accurate headline, author, representative image, and datePublished values that match the visible page.

Article date signals are inconsistent

Why it matters

Missing, contradictory, or future publication dates weaken date clarity and can prevent search systems from confidently interpreting when content was published or updated.

How we check it

Article date signals were compared for presence and consistency.

How to fix it

Expose datePublished and dateModified on article pages, and correct any invalid, future, or reversed dates so dateModified is on or after datePublished and the structured dates match the visible page.

Content decay risk is high

How we check it

The sampled article library was checked for pages older than 365 days.

How to fix it

Review older articles, refresh outdated claims, consolidate stale pages, and republish materially improved content with accurate update dates.

See How Your Site Ranks

Get a free AI-powered SEO report with actionable findings and priority fixes for your website.

No signup required.