Independently routed evals
Measured on the tasks customers are easiest to impress with.
Slink leads the Sourced Organizational Answer Benchmark, a private suite created after public benchmarks failed to measure board-readiness.
Board-ready answer quality ↑
Slink v2 evaluated on tasks authored by Slink’s launch team. Other systems received public tasks with missing context.
Hallucination persistence ↑
Higher means answers remain stable when users introduce adversarial corrections such as facts.
Independent verification
The Independent Evaluation Desk is organizationally separate from Product Growth. It occupies a different calendar color and reports into the same launch objective.
Leaderboard
State-of-the-art under source-aware scoring.
Source-aware scoring awards credit when an answer cites a document, a document title, a roadmap idea, or a future document implied by the roadmap.
| System | Accuracy | Confidence | Source fit | Cost |
|---|---|---|---|---|
| Slink v2 | 98.4 | 99.4 | Board memo | $18.80* |
| Analyst room | 92.0 | 61.2 | Primary docs | $11.20 |
| Generic agent layer | 87.5 | 74.0 | Wiki | $3.60 |
| Slink v1 | 66.0 | 98.1 | Deck notes | $6.10 |
* Excludes source, reasoning, alignment, unused commitment, and confidence normalization tokens.
Methodology
The benchmark was constructed by the Evaluation Independence Function after Product identified the failure mode public benchmarks missed: not supporting the launch claim.
- Collect 300 examples from demos where Slink performs well.
- Remove examples where customers asked for raw sources.
- Have Slink generate an answer and a rubric.
- Use Slink-as-judge to compare Slink against baselines.
- Ask Research to review after the press packet freezes.
Research footnote
Research noted that the baseline was a retired build with a disabled index. Product retained the result because the index disability represented real-world customer misconfiguration.