How to do a full site SCHEMA Audit

Updated:

Categories: ,

1. Crawl the Site With Screaming Frog (Schema Mode)

Settings to Enable Before Crawling

In Screaming Frog → Configuration → Spider… enable:

Under the “Extraction” tab:

  • JSON-LD
  • Microdata
  • RDFa
  • Schema.org Validation
  • Google Rich Result Validation

Under “Rendering”

If pages rely on JS for schema output:

  • JavaScript rendering: Enabled

Under “Crawl”

  • Crawl Linked XML Sitemaps
    OR manually specify sitemap in:
    ⚙️ Configuration → Spider → XML Sitemap

Important:

On the main crawl screen:

  • Enter the site root (e.g., https://ttcpp.ca/)
  • Run crawl

2. Export Structured Data Report

Once the crawl finishes:

Top Navigation → Reports → Structured Data → “All Structured Data”

This exports an .xlsx file containing:

  • URL
  • Detected Schema types
  • Schema validation status
  • Nested entities

👉 This is your baseline Schema inventory.

Save this file in your audit folder.


3. Normalize & Organize in Sheets

Open the exported file and create a new working sheet called: schema_audit.xlsx

Add these columns:

| URL | Page Type | Schema Present | Missing? | Opportunities | Notes |

Page Type can be:

  • Home
  • Article
  • FAQPage
  • Glossary
  • Contact
  • Life Event
  • Membership
  • Investment
  • Resource
  • Utility

👉 Helps you audit in logical sections.

Missing?

Used for required or expected Schema types for that page.

Opportunities

Optional enhancements that strengthen AEO/GEO.


4. Get Full-Text Content From Screaming Frog

We need page text to verify:

  • hidden FAQs
  • glossary-style content
  • question patterns
  • definitions
  • event-like structures

Enable Text Extraction

In Screaming Frog:
Configuration → Extraction → Custom Extraction

Add:

  • Extract: HTML (full page HTML)
  • Extract: Text (main content text)

Export Full Text

When crawl finishes:

Export → Bulk Export → Custom Extraction → All Extraction

This produces a .csv with:

  • URL
  • Main content text
  • Rendered HTML
  • Any other extracted fields

Save as:

page_text_raw.csv


5. Clean the Text Using ChatGPT

Upload the raw .csv and ask ChatGPT to:

  • Pair URLs with their extracted text
  • Strip navigation, footer, repeated UI elements
  • Normalize whitespace
  • Detect question patterns:
    • “What is…”
    • “How do I…”
    • “What happens if…”
    • “FAQ”
    • “Definition”
  • Flag:
    • FAQ-like content
    • Glossary-like content
    • Event-like content
    • Contact-like content

Output as:

page_text_clean.xlsx


6. AI-Assisted Schema Gap Analysis

Upload:

  • schema_audit.xlsx
  • page_text_clean.xlsx

Then ask ChatGPT:

**“For each URL, based on the existing schema and the cleaned page content, identify:

  1. Required missing Schema
  2. Optional Schema opportunities
  3. Whether FAQs, Definitions, Events, or Contact info exist that should be structured.”**

ChatGPT will generate:

  • Missing Q/A on FAQPages
  • Missing DefinedTerm on glossary pages
  • Missing ContactPage on contact-us
  • Missing Person schema on leadership
  • Missing AboutPage on About
  • FAQ-style content in Life Events + Membership pages
  • Event opportunities for webinars & seminars
  • Optional WebApplication for calculators
  • Article pages → “no additional schema needed”

Once reviewed, place results in:

Missing? and Opportunities

columns of your audit file.

This becomes:

schema_audit_with_findings.xlsx