1. Crawl the Site With Screaming Frog (Schema Mode)
Settings to Enable Before Crawling
In Screaming Frog → Configuration → Spider… enable:
Under the “Extraction” tab:
- JSON-LD
- Microdata
- RDFa
- Schema.org Validation
- Google Rich Result Validation
Under “Rendering”
If pages rely on JS for schema output:
- JavaScript rendering: Enabled
Under “Crawl”
- Crawl Linked XML Sitemaps
OR manually specify sitemap in:
⚙️ Configuration → Spider → XML Sitemap
Important:
On the main crawl screen:
- Enter the site root (e.g.,
https://ttcpp.ca/) - Run crawl
2. Export Structured Data Report
Once the crawl finishes:
Top Navigation → Reports → Structured Data → “All Structured Data”
This exports an .xlsx file containing:
- URL
- Detected Schema types
- Schema validation status
- Nested entities
👉 This is your baseline Schema inventory.
Save this file in your audit folder.
3. Normalize & Organize in Sheets
Open the exported file and create a new working sheet called: schema_audit.xlsx
Add these columns:
| URL | Page Type | Schema Present | Missing? | Opportunities | Notes |
Page Type can be:
- Home
- Article
- FAQPage
- Glossary
- Contact
- Life Event
- Membership
- Investment
- Resource
- Utility
👉 Helps you audit in logical sections.
Missing?
Used for required or expected Schema types for that page.
Opportunities
Optional enhancements that strengthen AEO/GEO.
4. Get Full-Text Content From Screaming Frog
We need page text to verify:
- hidden FAQs
- glossary-style content
- question patterns
- definitions
- event-like structures
Enable Text Extraction
In Screaming Frog:
Configuration → Extraction → Custom Extraction
Add:
- Extract: HTML (full page HTML)
- Extract: Text (main content text)
Export Full Text
When crawl finishes:
Export → Bulk Export → Custom Extraction → All Extraction
This produces a .csv with:
- URL
- Main content text
- Rendered HTML
- Any other extracted fields
Save as:
page_text_raw.csv
5. Clean the Text Using ChatGPT
Upload the raw .csv and ask ChatGPT to:
- Pair URLs with their extracted text
- Strip navigation, footer, repeated UI elements
- Normalize whitespace
- Detect question patterns:
- “What is…”
- “How do I…”
- “What happens if…”
- “FAQ”
- “Definition”
- Flag:
- FAQ-like content
- Glossary-like content
- Event-like content
- Contact-like content
Output as:
page_text_clean.xlsx
6. AI-Assisted Schema Gap Analysis
Upload:
schema_audit.xlsxpage_text_clean.xlsx
Then ask ChatGPT:
**“For each URL, based on the existing schema and the cleaned page content, identify:
- Required missing Schema
- Optional Schema opportunities
- Whether FAQs, Definitions, Events, or Contact info exist that should be structured.”**
ChatGPT will generate:
- Missing Q/A on FAQPages
- Missing DefinedTerm on glossary pages
- Missing ContactPage on contact-us
- Missing Person schema on leadership
- Missing AboutPage on About
- FAQ-style content in Life Events + Membership pages
- Event opportunities for webinars & seminars
- Optional WebApplication for calculators
- Article pages → “no additional schema needed”
Once reviewed, place results in:
Missing? and Opportunities
columns of your audit file.
This becomes:
schema_audit_with_findings.xlsx
