# SEO Checker Tool - Accuracy Limitations
Last Updated: 2026-09-06
# Overview
This document provides transparency about the aviary tool's accuracy limitations, known issues, and areas where manual verification is recommended.
This tool is designed to identify potential SEO issues. Not all findings indicate actual problems, and the tool cannot catch every SEO issue. Always apply professional judgment when interpreting results.
# 1. Previously Fixed Issues
# 1.1 Previously Disabled Checks ✅ FIXED
The following checks were disabled in earlier versions but have been re-enabled:
| Check | Location | Status | Description |
|---|---|---|---|
| Response Code Validation | Technical Checker | ✅ Fixed | Now properly checks HTTP status codes (200, 404, 500, etc.) |
| Compression Detection | Technical Checker | ✅ Fixed | Detects gzip, brotli, and deflate compression |
| Security Headers | Security Checker | ✅ Fixed | Validates HSTS, X-Frame-Options, CSP, X-Content-Type-Options |
| Cache Headers | Core Web Vitals | ✅ Fixed | Checks Cache-Control, ETag, Expires headers |
Previous Behavior: These checks always returned passed: true even when issues existed.
Fix: The tool now captures the initial HTTP response during navigation and passes it to all checkers that need HTTP headers, eliminating the execution context destruction issue.
# 1.2 Image Format Parsing Bug ✅ FIXED
Previous Issue: Image format detection showed invalid formats like:
co/67x84/d2df5b/656f10co/1044x532/9ca3af/374151
These were from placeholder/data URLs that weren't properly filtered.
Fix: Enhanced image format extraction to:
- Skip data URLs and placeholders
- Properly parse file extensions from URLs
- Detect WebP, AVIF, and other modern formats
- Fallback to MIME type when extension unavailable
# 1.3 Hidden Text Detection Improvements ✅ PARTIALLY FIXED
Previous Issue: Legitimate content was flagged as "hidden text spam":
- Collapsed accordions
- Tab content
- Off-screen navigation
- Truncated text with "read more" buttons
Fix: Updated hidden text detection to:
- Ignore common UI patterns (accordions, tabs, modals) on direct parents
- Check for legitimate accessibility hiding (screen readers)
- Reduce false positives for content overflow
- Only flag truly suspicious hiding techniques
# 1.4 Trust-of-Analysis Bug Sweep (2026-09-06) ✅ FIXED
Found by empirically running the tool against real sites (books.toscrape.com, demo.vercel.store) and hand-verifying flagged results against the raw DOM, then sweeping the rest of the codebase for the same two bug shapes:
Duplicated detection logic that had drifted out of sync (the same underlying fact re-derived independently by two checks, with no shared source of truth):
| Facts | Checks involved | Was | Fix |
|---|---|---|---|
| Charset declaration | internationalization.ts: charset-utf8, unicode-support-utf8 |
unicode-support-utf8 only checked meta[charset], false-failing pages (confirmed on books.toscrape.com) that declare charset via the older `` form |
Both call shared/dom.ts's new getCharset() |
| Viewport directives | mobileUX.ts, uiElements.ts, pageQuality.ts |
mobileUX.ts failed on maximum-scale (zoom lock); uiElements.ts never checked for it at all, silently passing the same tag |
All three parse via shared/dom.ts's new parseViewportMeta() (each keeps its own pass/fail policy) |
| Open Graph tags | metaTags.ts (og-tags-configured), socialMedia.ts (open-graph-configured) |
Each hand-rolled its own querySelectorAll('meta[property^="og:"]') |
Both call shared/dom.ts's new extractOgTags() |
| HTTPS/protocol | ecommerce.ts, legalCompliance.ts |
Each duplicated window.location.protocol === 'https:' in a browser evaluate() |
Both now use shared/dom.ts's new isHttpsUrl(this.page.url()), matching security.ts's existing Node-side technique |
Check-id collisions (two checkers registering the same rule id with different pass/fail criteria — since a rule id becomes a result's name, this conflates two different verdicts under one label in any flat, cross-checker view of a report):
| Id | Checkers | Fix |
|---|---|---|
product-schema-complete |
schemaValidation.ts, ecommerce.ts |
Renamed ecommerce.ts's to ecommerce-product-schema-complete |
dom-content-loaded-acceptable |
performance.ts, coreWebVitals.ts |
Renamed coreWebVitals.ts's to cwv-dom-content-loaded-acceptable |
page-size-acceptable |
technical.ts (raw HTML size), coreWebVitals.ts (total page weight) |
Renamed coreWebVitals.ts's to cwv-page-size-acceptable |
resource-hints-present |
resourceOptimization.ts, coreWebVitals.ts |
Renamed coreWebVitals.ts's to cwv-resource-hints-present |
A new test (tests/unit/registry.test.ts) now asserts no two checkers registered on BaseChecker share a rule id, so this bug class can't reappear silently.
Sibling/descendant text-measurement bug (confirmed false negative): ecommerce.ts's checkProductDescription summed textContent.length only over the descendants of elements matched by [class*="description"]/[id*="description"]. On books.toscrape.com, the matched container (<div id="product_description">) holds only a heading — the actual paragraph is a DOM sibling, not a child — so the check measured 41 characters against an actual description of several hundred, and false-failed. shared/dom.ts's new resolveDescriptiveText() falls back to a container's siblings when its own text looks like a bare label.
analyzeScrollDepth's coordinate bug (heatmap.ts, previously documented below in §2.4): fixed by replacing the elementsFromPoint probe with a document-relative bounding-box bucketing approach — see §2.4 for what the bug was.
# 1.5 Retest Findings (2026-09-06) ✅ FIXED
A follow-up retest against a broader set of real sites (webscraper.io's e-commerce test catalog, en.wikipedia.org, plus re-running books.toscrape.com and demo.vercel.store) surfaced two more issues, one of them introduced by §1.4's own fix:
resolveDescriptiveText() over-padding a genuinely short description: the sibling-rescue fallback added in §1.4 correctly fixed the books.toscrape.com case, but on a real product card (webscraper.io) it also pulled in a price and title element that happened to be siblings of a genuinely short (99-character) description, padding the count to 166 and passing a description that should have failed by one character. Fixed by excluding siblings that are themselves a different structured product field (detected via itemprop or a price/title/name/sku/brand naming convention) from the fallback — it now only rescues text that was actually misplaced, not any nearby text.
A serialization hazard in the same fix, caught only by manually running the CLI: the first version of that exclusion logic used a nested helper function (const isOtherStructuredField = (el) => {...}) declared inside resolveDescriptiveText. Running the tool via npx tsx src/cli.ts (the dev-mode runner used throughout this project's own testing) crashed with ReferenceError: __name is not defined — tsx's esbuild-based transform wraps nested function declarations with a name-preservation helper call that isn't included when Playwright serializes just the outer function's source via .toString() for page.evaluate(). The crash was caught by the check's existing try/catch and degraded gracefully to "check skipped" rather than crashing the audit — so real-world impact was silent under-reporting, not a hard failure. This did not affect the actual published package: the production build (tsc, via npm run build:ts) and the Vitest test suite (a different esbuild configuration) both compile the nested closure as plain JS with no such wrapper, and neither was affected — confirmed by building and running the compiled dist/cli.js with the buggy version in place. Fixed by inlining the check directly rather than declaring a nested named function, matching this file's own top-of-file rule that every function passed to page.evaluate() must be fully self-contained. A real-browser end-to-end test (tests/e2e/seoChecker.e2e.test.ts's "resolves product-description-present against a real page without crashing") now exercises this function against an actual Playwright page rather than the mock DOM every other unit test uses — closing the specific blind spot where mock-DOM tests can't validate that a function actually survives Playwright's serialization boundary. That test does not reproduce the tsx-specific wrapper (Vitest doesn't inject it either), so the real protection against this exact hazard going forward is the "no nested closures" code pattern, not the test.
# 2. Inherent Limitations & Code Bugs (Heuristic-Based)
These checks use statistical models or heuristics that cannot be 100% accurate, or contain specific implementation bugs:
# 2.1 Readability Scores (~85% accurate)
Check: Content Readability (Flesch-Kincaid, Gunning Fog)
Limitation:
- Based on syllable counting and sentence length
- Cannot understand context or domain complexity
- Medical/legal content will score poorly despite being appropriate
- Creative writing may score unexpectedly
Recommendation: Use as a guideline, not absolute rule. Consider your target audience's education level.
# 2.2 Spam Detection (~60% accurate)
Check: Spam Patterns, Keyword Stuffing, Hidden Text
Limitation & Code Bugs:
- Pattern-based detection has false positives
- Cannot understand intent
- Shallow DOM Hidden Text check bug: In
spamDetection.ts, theisLegitimateHidden()function only evaluates the hidden element itself and its direct parent (el.parentElement) for accordion or collapse framework classes. In Tailwind and Bootstrap components, interactive container classes (such as.collapseor.accordion) are often located on higher ancestors. Because the checker does not traverse up the DOM tree, it flags these legitimate hidden elements as potential hidden text spam. - Keyword density thresholds are heuristic-based
- Industry-specific terminology may be flagged as repetitive
Known False Positives:
- Product descriptions with natural keyword repetition
- Legal disclaimers with repeated terms
- Multi-language content
- Lists of similar items (product catalogs)
Recommendation: Manually review flagged items. High spam scores (>70%) are more reliable.
# 2.3 Content Quality Assessment (~70% accurate)
Check: Content Depth, Uniqueness, Structure
Limitation & Code Bugs:
- Cannot judge factual accuracy
- Cannot assess expertise or authority
- Regulatory Auditing Omission Gap: In
legalCompliance.ts, the checks for GDPR and CCPA returnpassed: trueif their respective compliance terms are missing. This means if a site completely lacks a privacy policy or regulatory statements, the checker still passes instead of warning or failing.
What It Can Detect:
- ✅ Thin content (word count)
- ✅ Poor structure (headings)
- ✅ Missing key elements
What It Cannot Detect:
- ❌ Plagiarism from other sites
- ❌ Factual errors
- ❌ Content relevance to search intent
- ❌ E-A-T signals (Expertise, Authority, Trust)
# 2.4 Mobile Usability & Heatmaps (~75% accurate)
Check: Tap Target Size, Viewport Configuration, Scroll Depth
Limitation & Code Bugs:
- 44px tap target rule is a guideline (WCAG 2.5.5)
- Viewport simulation vs. actual device behavior
- Scroll Depth Coordinate Bug ✅ FIXED (see §1.4):
heatmap.ts's scroll depth content density checker used to pass the document-relative vertical offsetyPositionintodocument.elementsFromPoint(), which expects viewport-relative client coordinates -- since the audit never actually scrolls the page, anyyPositionbeyond one viewport height returned an empty array, zeroing the density score for nearly every depth band on a typical page. Replaced with a bounding-box bucketing approach that doesn't depend on the page having scrolled there.
Recommendation: Test on real devices for critical pages.
# 2.4a Heatmap & Click Prediction (no ground truth available)
Check: Click Heatmap, Attention Zones (heatmap.ts's generateClickHeatmap and analyzeAttentionZones)
Unlike every other check in this document, these aren't measuring a DOM fact that can be right or wrong -- they assign ad-hoc weighted scores (element type, size, position, background color) modeling where a real user would click or look. There is no ground truth available from a static crawl: real click/attention data comes from recorded user sessions (Hotjar, Microsoft Clarity, GA4 scroll-depth), which this tool has no access to. heatmap.ts is the only checker in the codebase built this way.
What this means in practice:
- Messages are worded "predicted"/"estimated" deliberately, not decoratively -- they should never be read as measured facts the way, say, an HTTPS check result is.
- What can be validated without ground truth: the relative ranking makes sense (a colored above-fold CTA should outscore a buried below-fold link) and the scoring doesn't silently drift (a regression that swapped two weight constants should be caught by a test, not ship silently).
tests/unit/heatmap.test.tshas rank-plausibility and exact-score-pinning tests for this. - What can't be validated: whether the absolute scores correlate with real user behavior on any given site. That requires correlating against actual analytics on a live, operated site -- out of scope for a static audit tool.
Recommendation: Treat heatmap scores as a heuristic prioritization aid (which elements should draw attention, per visual-hierarchy best practice), not as a substitute for real user analytics.
# 2.5 Storage & Cookie Consent Verification (~65% accurate)
Check: Cookie Consent Banner, Cookie Policy Link (legalCompliance.ts: cookie-consent-present, cookie-policy-linked)
Limitation & Code Bugs:
- No Runtime Cookie or Storage Inspection: The tool does not inspect actual browser cookies (
page.context().cookies(),document.cookie), HTTPSet-Cookieresponse headers, or client storage APIs (localStorage,sessionStorage,IndexedDB). - Superficial DOM Selector Matching:
checkCookieConsentevaluates whether DOM elements match[class*="cookie"],[id*="cookie"],[class*="consent"], or buttons containing keywords like "accept" or "consent". - False Positives on Cookieless Sites: A static site that sets zero cookies and stores no personal data is flagged as failing if it lacks a cookie consent banner, even though cookieless sites legally require no consent mechanism under GDPR and the ePrivacy Directive.
- False Negatives on Non-Compliant Sites: A site with a non-functional or decorative consent banner that drops tracking cookies prior to user opt-in will still pass because the checker only tests for the presence of DOM elements, not actual consent gating.
Manually inspect DevTools Application/Storage tabs and network headers to verify actual cookie behavior and local storage persistence.
# 3. Client-Side Architectural Limitations
These limitations stem from the tool running in a browser context:
# 3.1 Network Timing Variability
Limitation:
- Performance metrics vary per run
- Network conditions affect results
- Geographic location matters
Recommendation:
- Run multiple checks and average results
- Use dedicated performance tools (Lighthouse, WebPageTest) for detailed analysis
# 3.2 JavaScript Execution Required
Limitation:
- Only sees what JavaScript renders
- Cannot test "JavaScript disabled" experience
- May miss noscript content
# 3.3 Cannot Verify Actual Indexing
Limitation:
- Tool checks if page is indexable, not if it's indexed
- Cannot verify Google's actual index status
# 3.4 Core Web Vitals Measurement Approximations
The Core Web Vitals category (coreWebVitals) measures real LCP, CLS, FCP, and TTFB via the standard web-vitals library, injected into the page before navigation so its observers can see load-time entries. Two disclosed approximations follow directly from running as an unattended, single-shot audit rather than a real browser session:
- Latest-value, not final-value.
web-vitalsnormally reports a metric's final value when the page is navigated away from or the tab is hidden — neither ever happens here, since the audit closes the browser outright. Metrics are instead collected withreportAllChanges: trueand read at the same point every other checker reads the page (afternetworkidleplus a stability wait). For a page that has finished loading, this is normally the same value a real session would report, but it isn't guaranteed down to the millisecond. - No real INP — Total Blocking Time substitutes. INP (Interaction to Next Paint) requires a real user interaction (click, tap, keypress) to measure, and this audit never interacts with the page — there's no honest way to synthesize one. Rather than fabricate an interaction to claim an "INP" number, the
total-blocking-time-acceptablecheck reports Total Blocking Time (summedlongtaskentries over the 50ms threshold) as a disclosed lab proxy for interactivity — the same substitution Lighthouse makes, and for the same reason.
# 4. Missing Production Features & Hidden Behaviors
Not yet implemented, or undocumented CLI behaviors worth knowing about:
- ❌ Parallel URL checking (checking multiple URLs in one run)
- ❌ Caching mechanisms (reusing results from previous runs)
- ❌ Lighthouse integration (Google's official tool)
- ❌ Google Search Console API integration
- ❌ Historical data tracking and trend analysis
# 4.1 Missing OpenAI Provider in Rust Engine
.env.examplelistsopenaias a valid value forAVIARY_LLM_PROVIDER, butengine/src/semantic/factory.rsonly implementsollamaandstub. Setting the provider toopenaisilently falls back to theStubAnalyzerrather than erroring.
# 4.2 Prometheus Metrics Server Starts on Import
Importing the CLI registers and starts a Prometheus metrics server (prom-client), configurable via AVIARY_METRICS_PORT (default 9090). It starts silently in the background as a side effect of import rather than an explicit opt-in, which can surprise anything embedding src/cli.ts as a library and may conflict with another local service already on that port.
# 5. Known False Positives by Category
# 5.1 Meta Tags & SEO Basics (95% accurate)
Rare False Positives:
- Brand information in non-standard meta tags
- Alternative meta tag implementations (custom CMS)
- Structured data in non-JSON-LD formats
# 5.2 Structured Data (90% accurate)
Known Issues:
- May flag valid but uncommon schema types
- Nested schema validation can be overly strict
- Custom schema extensions may not validate
# 5.3 Performance Metrics (85% accurate)
Known Issues:
- Network timing varies ±20% per run; a single-run measurement may not represent typical performance
- Doesn't account for CDN edge caching
- First visit vs. cached visit differences
- Cannot detect server-side rendering optimizations or HTTP/2 push resources
Recommendation: Run multiple checks and cross-reference with a dedicated tool (Lighthouse, PageSpeed Insights, WebPageTest) for production analysis.
# 5.4 Accessibility (80% accurate)
Known Issues:
- Color contrast calculation doesn't account for gradients
- ARIA validation may flag valid custom implementations
- Cannot evaluate alt text quality, only presence
- Cannot test keyboard navigation flows or verify actual screen reader compatibility
- May miss dynamically loaded content
Recommendation: Supplement with manual testing using an actual screen reader and keyboard-only navigation.
# 5.5 Image Optimization (75% accurate)
Known Issues:
- Cannot verify actual compression quality
- CDN Format Detection Limit: CDNs like Cloudinary are not automatically recognized as
'dynamic'in the TScdnPatternsarray (src/checkers/advancedImages.ts). Only common placeholder sites (e.g.placehold.co,dummyimage.com) are correctly categorized as dynamic placeholders. Other CDNs fall back to raw file extensions or are marked as'unknown'.
# 5.6 Spam Detection (60% accurate)
High False Positive Rate:
- Product descriptions with natural keyword density
- Technical documentation with repeated terms
- Legal disclaimers
- Interactive elements inside nested container components (accordion, collapse, modal) due to shallow DOM traversal limit.
# 5.7 Storage & Cookie Consent (65% accurate)
Known Issues:
- Cookieless sites flagged for missing consent banners (consent not required if no cookies/trackers used).
- Custom consent management platforms (CMPs) rendered via shadow DOM or canvas unflagged.
- Cannot inspect actual cookies,
Set-CookieHTTP headers, or client-side storage keys (localStorage,sessionStorage,IndexedDB).
# 6. Accuracy Estimates by Check Type
| Check Category | Accuracy | Confidence Level | Notes |
|---|---|---|---|
| Meta Tags | ~95% | High | Straightforward DOM parsing |
| Heading Structure | ~95% | High | Clear hierarchy rules |
| HTTPS/Security | ~90% | High | Binary checks (present/absent) |
| Structured Data | ~90% | High | Schema.org validation |
| Response Codes | ~90% | High | HTTP standard compliance |
| Compression | ~90% | High | Header presence check |
| Performance Metrics | ~85% | Medium | Network variability |
| Accessibility | ~80% | Medium | Complex WCAG rules |
| Mobile Usability | ~75% | Medium | Viewport simulation vs real devices; tap target heuristic (WCAG 2.5.5) |
| Image Optimization | ~75% | Medium | CDN detection limitations |
| Content Quality | ~70% | Medium | Regulatory compliance gap |
| Storage & Cookie Consent | ~65% | Low | Superficial DOM keyword matching; no runtime cookie or storage inspection |
| Spam Detection | ~60% | Low | Shallow DOM check, false positives |
| Readability | ~70% | Low | Statistical estimation |
| Heatmap & Click Prediction | ~50% | Low | Heuristic weighting model; no live user session ground truth |
# 7. Best Practices for Using This Tool
# 7.1 Interpretation Guidelines
- Errors (Red): Address these - likely real issues
- Warnings (Yellow): Review manually - may be false positives
- Info (Blue): Suggestions - consider for optimization
# 7.2 Verification Workflow
For critical findings:
- ✅ Run check 2-3 times to confirm consistency
- ✅ Cross-reference with official tools (Google Search Console, Rich Results Test)
- ✅ Manual inspection in browser DevTools
- ✅ Test on real devices (mobile checks)
# 7.3 Priority-Based Actions
High Priority (Fix Immediately):
- ✅ Missing title/meta description
- ✅ Broken HTTPS/mixed content
- ✅ 404/500 response codes
- ✅ Mobile viewport not set
- ✅ No robots.txt
Medium Priority (Review & Fix):
- ⚠️ Missing structured data
- ⚠️ Slow performance metrics
- ⚠️ Accessibility violations
- ⚠️ Missing alt attributes
- ⚠️ Broken links
Low Priority (Consider Optimization):
- ℹ️ Image format suggestions
- ℹ️ Readability improvements
- ℹ️ Additional schema markup
- ℹ️ Content length recommendations
# 8. Reporting Issues
If you encounter false positives or inaccurate checks:
- Verify: Is this actually incorrect?
- Report: Create GitHub issue with tested URL, failed check, and expected behavior.
GitHub Issues: https://github.com/Ru1vly/Aviary/issues
# 9. Conclusion
The aviary tool is most accurate for:
- ✅ Technical SEO fundamentals (meta tags, headers)
- ✅ Structural issues (headings, links)
- ✅ Basic accessibility
- ✅ HTTPS/security checks
- ✅ Structured data validation
Use with caution for:
- ⚠️ Spam detection (shallow DOM checking)
- ⚠️ Content quality assessment (subjective and regulatory gaps)
- ⚠️ Performance metrics (network variability)
- ⚠️ Readability scores (domain-dependent)
- ⚠️ Heatmaps & Click prediction (predictive weighting without real user session ground truth)
- ⚠️ Storage & Cookie consent (superficial DOM keyword matching; no runtime cookie or storage inspection)
This document will be updated as the tool evolves and accuracy improves.