AI Page Discovery — Test Cases
Unit Tests
| Case | Expected Result |
|---|---|
| Homepage with sitemap and duplicate links | Deduplicated same-origin HTML pages |
| Fragments and tracking query parameters | Removed; meaningful parameters retained |
| Private DNS address or external redirect | Rejected before connection |
| Robots disallow or robots retrieval failure | Pages not fetched |
| Action link, video, image, or download | Excluded |
| Page, depth, sitemap, or frontier cap | Work stays within limits; page, sitemap, and frontier truncation reported |
| Checkpoint followed by continuation | Remaining URLs processed without losing progress |
| Completed message replay | Existing result retained without refetching |
| S3 or SQS failure | Record returned for retry |
| Project switch during plugin polling | Previous project's response ignored |
| Select preview URLs | Manual import field populated; import requires user action |
Direct .xml sitemap input |
Sitemap URLs listed without fetching robots.txt or page HTML |
| Sibling URLs | Grouped by their exact parent path regardless of HTML structure |
| Plugin discovery preview | URLs shown as a nested path tree with no default selection |
Local Verification
Run uv run pytest apps/page-discovery/__tests__/. Use mocked repository and controlled HTTP fixtures; do not crawl customer websites during automated tests. Run backend batch tests and plugin UI tests separately. AWS deployment is not part of local verification.