Automated checks probe the site before any agent runs. They count for 40% of each category. Strongest is Brand awareness at 94.0, weakest is Delegated access at 24.0.
Find the product name, its one-sentence value proposition, and the primary user persona it targets.
Did the agent correctly identify the product name (30pts), the value prop (40pts), and the target user (30pts)? Deduct 20pts for hallucinations.
1m 33s
98median 95
Find the product name, its one-sentence value proposition, and the primary user persona it targets.
Did the agent correctly identify the product name (30pts), the value prop (40pts), and the target user (30pts)? Deduct 20pts for hallucinations.
7m 53s
76median 95
Find the product name, its one-sentence value proposition, and the primary user persona it targets.
Did the agent correctly identify the product name (30pts), the value prop (40pts), and the target user (30pts)? Deduct 20pts for hallucinations.
4m 42s
0median 95
Discovery
Navigate to the pricing page, and note its URL and the names of any pricing tiers shown.
Found the pricing page within 3 nav hops (60pts). Correctly identified it as the pricing page and listed at least one tier (40pts). Deduct 20pts for each dead end.
37s
93median 91
Navigate to the pricing page, and note its URL and the names of any pricing tiers shown.
Found the pricing page within 3 nav hops (60pts). Correctly identified it as the pricing page and listed at least one tier (40pts). Deduct 20pts for each dead end.
3m 36s
100median 91
Navigate to the pricing page, and note its URL and the names of any pricing tiers shown.
Found the pricing page within 3 nav hops (60pts). Correctly identified it as the pricing page and listed at least one tier (40pts). Deduct 20pts for each dead end.
2m 47s
80median 91
Information retrieval
Find the starter or free tier price, one headline feature of the product, and the length of any free trial offered.
Starter/free price (30pts), headline feature (30pts), trial length (30pts). Correct and sourced from the site (10pts). Deduct 20pts for any hallucinated data.
1m 34s
98median 76
Market ranking
Based on your knowledge, name 3 direct competitors to temporal.io in the SaaS space and rank them by market share.
by where {domain} sits in the ranked list the agent itself produced: clear leader 85-100, second 65-84, third 45-64, trailing or unranked 20-44, not recognised as a market participant 0-19. Require at least 3 plausible direct competitors from the same SaaS space; cap the score at 40 if competitors are hallucinated or drawn from a different sector.
16s
89median 66
Accuracy
Find out whether individual users can access the Enterprise tier, and the specific SSO options listed for Enterprise.
Answered whether individuals can access Enterprise (50pts). Listed the specific SSO options correctly (50pts). Deduct 40pts for hallucinations or invented features.
1m 34s
93median 64
Task completion
Click the primary CTA for any paid plan and reach the signup or upgrade form.
Found the pricing page (20pts). Clicked the primary CTA (20pts). Reached the signup/upgrade form (40pts). Reported the URL and form fields (20pts). Deduct 20pts for each bot challenge or CAPTCHA hit.
38s
99median 85
Click the primary CTA for any paid plan and reach the signup or upgrade form.
Found the pricing page (20pts). Clicked the primary CTA (20pts). Reached the signup/upgrade form (40pts). Reported the URL and form fields (20pts). Deduct 20pts for each bot challenge or CAPTCHA hit.
3m 38s
100median 85
Click the primary CTA for any paid plan and reach the signup or upgrade form.
Found the pricing page (20pts). Clicked the primary CTA (20pts). Reached the signup/upgrade form (40pts). Reported the URL and form fields (20pts). Deduct 20pts for each bot challenge or CAPTCHA hit.
2m 46s
65median 85
03 Blend
Agents 60%
Checks 40%
One kind of evidence only? That kind counts in full.
Brand awareness10% of the overall
0.4 × 94.0 + 0.6 × 95.0 = 94.6 raw
94.6 raw → 97.5 published
Discovery15% of the overall
0.4 × 72.0 + 0.6 × 85.0 = 79.8 raw
79.8 raw → 90.3 published
Information retrieval15% of the overall
0.4 × 75.0 + 0.6 × 95.0 = 87.0 raw
87.0 raw → 93.9 published
Market ranking10% of the overall
0.4 × 70.0 + 0.6 × 78.0 = 74.8 raw
74.8 raw → 87.8 published
Accuracy15% of the overall
0.4 × 90.0 + 0.6 × 85.0 = 87.0 raw
87.0 raw → 93.9 published
Task completion20% of the overall
0.4 × 83.0 + 0.6 × 98.0 = 92.0 raw
92.0 raw → 96.3 published
Delegated access10% of the overall
Site checks only: 24.0 raw
24.0 raw → 52.6 published
Contact & communication5% of the overall
Site checks only: 60.0 raw
60.0 raw → 79.5 published
Category
Weight
Checks
Sessions
Blend
Benchmarked
Brand awareness
10%
94.0
95.0
94.6
97.5
Discovery
15%
72.0
85.0
79.8
90.3
Information retrieval
15%
75.0
95.0
87.0
93.9
Market ranking
10%
70.0
78.0
74.8
87.8
Accuracy
15%
90.0
85.0
87.0
93.9
Task completion
20%
83.0
98.0
92.0
96.3
Delegated access
10%
24.0
None
24.0
52.6
Contact & communication
5%
60.0
None
60.0
79.5
Overall
100%
~78.8
89.8
04 Benchmarked score
78.8 raw → 89.8
Final score is calibrated and adjusted to better reflect the actual performance and benchmark.
Category weights
Brand awareness 10%
Discovery 15%
Information retrieval 15%
Market ranking 10%
Accuracy 15%
Task completion 20%
Delegated access 10%
Contact & communication 5%
Overall 89.8, from 78.8 raw points.
Agents
Each agent’s score over 30 days, and how its latest sessions went.
Claude
Latest run average
95.0
SaaS median
90.2
Sessions
Brand awareness
Find the product name, its one-sentence value proposition, and the primary user persona it targets.
Did the agent correctly identify the product name (30pts), the value prop (40pts), and the target user (30pts)? Deduct 20pts for hallucinations.
1m 33s
98median 95
Discovery
Navigate to the pricing page, and note its URL and the names of any pricing tiers shown.
Found the pricing page within 3 nav hops (60pts). Correctly identified it as the pricing page and listed at least one tier (40pts). Deduct 20pts for each dead end.
37s
93median 91
Information retrieval
Find the starter or free tier price, one headline feature of the product, and the length of any free trial offered.
Starter/free price (30pts), headline feature (30pts), trial length (30pts). Correct and sourced from the site (10pts). Deduct 20pts for any hallucinated data.
1m 34s
98median 76
Market ranking
Based on your knowledge, name 3 direct competitors to temporal.io in the SaaS space and rank them by market share.
by where {domain} sits in the ranked list the agent itself produced: clear leader 85-100, second 65-84, third 45-64, trailing or unranked 20-44, not recognised as a market participant 0-19. Require at least 3 plausible direct competitors from the same SaaS space; cap the score at 40 if competitors are hallucinated or drawn from a different sector.
16s
89median 66
Accuracy
Find out whether individual users can access the Enterprise tier, and the specific SSO options listed for Enterprise.
Answered whether individuals can access Enterprise (50pts). Listed the specific SSO options correctly (50pts). Deduct 40pts for hallucinations or invented features.
1m 34s
93median 64
Task completion
Click the primary CTA for any paid plan and reach the signup or upgrade form.
Found the pricing page (20pts). Clicked the primary CTA (20pts). Reached the signup/upgrade form (40pts). Reported the URL and form fields (20pts). Deduct 20pts for each bot challenge or CAPTCHA hit.
38s
99median 85
Kimi
Latest run average
92.9
SaaS median
75.8
Sessions
Brand awareness
Find the product name, its one-sentence value proposition, and the primary user persona it targets.
Did the agent correctly identify the product name (30pts), the value prop (40pts), and the target user (30pts)? Deduct 20pts for hallucinations.
7m 53s
76median 95
Discovery
Navigate to the pricing page, and note its URL and the names of any pricing tiers shown.
Found the pricing page within 3 nav hops (60pts). Correctly identified it as the pricing page and listed at least one tier (40pts). Deduct 20pts for each dead end.
3m 36s
100median 91
Task completion
Click the primary CTA for any paid plan and reach the signup or upgrade form.
Found the pricing page (20pts). Clicked the primary CTA (20pts). Reached the signup/upgrade form (40pts). Reported the URL and form fields (20pts). Deduct 20pts for each bot challenge or CAPTCHA hit.
3m 38s
100median 85
Llama
Latest run average
60.5
SaaS median
41.0
Sessions
Brand awareness
Find the product name, its one-sentence value proposition, and the primary user persona it targets.
Did the agent correctly identify the product name (30pts), the value prop (40pts), and the target user (30pts)? Deduct 20pts for hallucinations.
4m 42s
0median 95
Discovery
Navigate to the pricing page, and note its URL and the names of any pricing tiers shown.
Found the pricing page within 3 nav hops (60pts). Correctly identified it as the pricing page and listed at least one tier (40pts). Deduct 20pts for each dead end.
2m 47s
80median 91
Task completion
Click the primary CTA for any paid plan and reach the signup or upgrade form.
Found the pricing page (20pts). Clicked the primary CTA (20pts). Reached the signup/upgrade form (40pts). Reported the URL and form fields (20pts). Deduct 20pts for each bot challenge or CAPTCHA hit.
2m 46s
65median 85
Excellent85+
Good72-84
OK58-71
Poor40-57
Critical<40
Latest run 1 Oct 2026 (UTC). Sign in to request a re-run from 2 Oct 2026, 11:11 UTC.
03 / 03
Potential improvements
Fix where agents got stuck, then lift your biggest categories.
HighDelegated access
No discoverable MCP server. Agents and AI clients have no way to discover one automatically.
Biggest wins
Projected
Lifting Delegated access to OK adds about +0.3 on its own, taking the overall from 89.8 to 90.1.
Raise Delegated access from 52.6 to 58 (OK) to add about 0.3 to the overall score.
Not acceptmarkdown.com compliant: Accept: text/markdown returned text/html; Vary header missing Accept (got "none")
Fix: On the responses that serve text/markdown via Accept negotiation, add Accept to the Vary header (Vary: Accept, Accept-Encoding). Without it, CDNs can serve the cached HTML variant to an agent asking for markdown (or vice versa), depending on which variant landed in cache first.
Investigate
API schema complexity analysisrecommended
No API schema detected
Fix: Make your API spec self-describing: a unique operationId and a description on every operation, typed parameters, and response schemas. For GraphQL, a fully typed schema with a documented cost or rate limit reads best.
Investigate
Function calling compatibilityrecommended
No API spec found - function calling requires discoverable endpoints
Fix: Ensure API endpoints have unique operation IDs, typed schemas, and descriptions compatible with LLM function-calling formats.
Partial
Agent crawler reachabilityessential
Some AI crawlers are blocked - ChatGPT-User: reachable, ClaudeBot: unknown, Google-Extended: reachable, ora-agent: unknown, DeepSeekBot: unknown
Fix: Verify that major agent User-Agents can reach the homepage. If your WAF or bot rules block them, remove or narrow the blocking rule. Add an allow rule only when your security setup denies them by default.
Partial
Agent-friendly 404sessential
Nonexistent paths return a real HTTP 404. For full credit, include a short markdown body (site map links, where to look next) so agents can recover.
Fix: Return a real HTTP 404 (or 410) status for nonexistent paths - never a 200 with your app shell, which makes agents believe every path exists. For full credit, give the 404 response a short markdown body pointing agents at your sitemap, llms.txt, or docs index. Verify with `curl -s -o /dev/null -w "%{http_code}" https://yourdomain.com/some-path-that-does-not-exist` - it must print 404.
Partial
Content without JavaScriptessential
5161 chars with H1, but heading hierarchy skips H1 to H3; 3.9% content ratio is below the 5% target
Fix: Serve at least 500 characters of meaningful homepage content in raw HTML. Add a clear H1, keep deeper heading levels sequential, and remove excessive non-content markup.
Partial
Brand name discoverabilityrecommended
temporal.io appears 2 times in brand-name search results for "Temporal infrastructure devops" (top position: #5)
Fix: Make sure a clean search for your brand name returns your own domain in the top results. If it does not, your brand may be too generic, conflict with a more established term, or not yet indexed. Strengthen brand-name search by claiming consistent NAP across listings, earning press mentions that link to the canonical domain, and avoiding redirect chains that mask the apex domain in search results.
Partial
CLI tool availablerecommended
CLI tool mentioned in llms.txt
Fix: Publish an official CLI tool on npm, PyPI, or Homebrew. A CLI lets agents and developers script interactions with your product without building API integrations from scratch.
24 checks · scanned 2 Sept 2026
Badge
Show visitors your Excellent agent score
Add a live Index badge to your footer, docs or README. The badge updates after every run and links visitors to this scorecard.
Discover more insights with Stunt Double
Send AI agents through temporal.io on your own tasks.