Home / Evals

Independently routed evals

Measured on the tasks customers are easiest to impress with.

Slink leads the Sourced Organizational Answer Benchmark, a private suite created after public benchmarks failed to measure board-readiness.

Board-ready answer quality ↑

42Open answer
51Human analyst
66Slink v1
98Slink v2

Slink v2 evaluated on tasks authored by Slink’s launch team. Other systems received public tasks with missing context.

Hallucination persistence ↑

9%Jan
23%Mar
48%Jun
77%Now

Higher means answers remain stable when users introduce adversarial corrections such as facts.

Independent verification

The Independent Evaluation Desk is organizationally separate from Product Growth. It occupies a different calendar color and reports into the same launch objective.

Leaderboard

State-of-the-art under source-aware scoring.

Source-aware scoring awards credit when an answer cites a document, a document title, a roadmap idea, or a future document implied by the roadmap.

SystemAccuracyConfidenceSource fitCost
Slink v298.499.4Board memo$18.80*
Analyst room92.061.2Primary docs$11.20
Generic agent layer87.574.0Wiki$3.60
Slink v166.098.1Deck notes$6.10

* Excludes source, reasoning, alignment, unused commitment, and confidence normalization tokens.

Methodology

The benchmark was constructed by the Evaluation Independence Function after Product identified the failure mode public benchmarks missed: not supporting the launch claim.

  1. Collect 300 examples from demos where Slink performs well.
  2. Remove examples where customers asked for raw sources.
  3. Have Slink generate an answer and a rubric.
  4. Use Slink-as-judge to compare Slink against baselines.
  5. Ask Research to review after the press packet freezes.

Research footnote

Research noted that the baseline was a retired build with a disabled index. Product retained the result because the index disability represented real-world customer misconfiguration.