I’m building Signalpost at Builderr: a test for whether a company-research agent can produce a result you can actually audit.
Give it a Norwegian organisation number. The agent returns public company facts with the source URL and retrieval date for each claim, and marks uncertainty instead of filling gaps. The starter runs one saved company first, so you can inspect a wrong match or weak citation before you scale a crawler.
The question I’m working on: if several crawlers agree, how do you tell whether they all missed the same source? I’m treating agreement as a clue, not recall, and looking for a practical test-harness pattern that makes the shared blind spot visible.
Challenge and starter: https://builderr.ai/challenges/signalpost?utm_source=slop&utm_medium=community&utm_campaign=builder_acquisition_20260926&utm_content=shared_blind_spots_build_v1
Disclosure: I run Builderr.
Live Demo