<- all tokdocs

Clarity Harness Traces Every Requirement Back to a Transcript, and Its Own Demo Shows 44 Verdicts Waiting

Watch on TikTok

View on TikTok ->

Robert Ta, co-founder of Epistemic Me, spends 91 seconds screen-recording an internal tool called the Clarity Harness. Four agents produce user stories from a recorded conversation, and every requirement carries a link back to the sentence a human actually said. The interesting part is not the agent tree, it is the review surface: a panel where a product manager reads the requirement and the source quote side by side, then approves, annotates, or rejects it with a keystroke. The screen shows a left rail labeled Clarity Harness, a Mission control pane with filter chips for pm-story-writer, clarity-orchestrator-agent, pm-agent, clarity-interview-agent, pm-requirements-synthesizer, and spec-align, and a judgment view with five tabs (Verdict, Chain, Delta, Trace, Artifact). The video never says that the harness is unavailable. As of 2026-09-29, heyclarity.dev is a waitlist page, the repo is not public, and the Epistemic-Me GitHub organization contains no harness repository.

What is actually on the screen

The demo opens on a scrolled table of archetypes: Prototyper, Builder, Sweeper, Grower, Maintainer. Each row carries three capability chips (Prototyper gets "parallel prototyping", "idea generation", "kill decisions") and a provenance line reading "drawn from Product manager · Designer · Engineer".

Mission control follows. The left rail shows a session named astek-suzie, "6 agents · 56 runs", and "runs today: 0". The main pane groups runs by agent. Clarity Orchestrator shows 18 runs, 1 reviewed, 15 verdicts waiting, v0.1.0, idle. Spec Align shows 22 runs, 10 reviewed, 12 verdicts waiting. PM Requirements Synthesizer shows 18 runs, 1 reviewed, 9 verdicts waiting. Individual rows are real-looking work items: "Lynn: Dayforce services transformation, WFM extraction agent in two months" and "CLA-320 design review, 20 world-class E2E refinements, 10 scoped for M3". A footnote reads "8 artifacts labeled auto_approved_dogfood this session · 7-day revert · audit-sampled".

Clicking through opens the judgment view, headed "JUDGMENT · DECIDE THE FOCUSED UNIT, verdicts train the next run". The Verdict tab shows a block labeled "SOURCE EVIDENCE, THE HUMAN'S WORDS" containing a quote from an email thread titled "FW: Epistemic Me and Services Transformation, Robert × Lynn", plus a "widen to the source" control and an "AGENT'S RATIONALE" section. The Chain tab shows the same quote tagged boundary: customer_requirement_evidence. The Artifact tab shows the drafted requirement with Approve (A), Annotate (E), and Reject (R) shortcuts. The annotate panel is labeled "your words become the training signal" and offers preset chips: too broad, not testable, misread what was said, wrong source cited, too confident, split into two.

The right-hand column is the trace, headed "FLOW · WHAT ACTUALLY HAPPENED, click a step to drive the fold", with a legend for model call, tool call, decision, path taken, not taken, state carried, and dropped. Step 01 is a tool call that ingests the source packet and resolves one source. Step 02 is a model call that extracts requirements and proposes four new ones with IDs like DAY-REQ-WFM-IR-001. Edge labels between steps read "route by boundary" and "coherence: dedupe + boundary".

Requirement traceability is a thirty-year-old problem, and this attacks the hard half

Gotel and Finkelstein studied this in 1994 at the first IEEE International Conference on Requirements Engineering, drawing on interviews with over 100 practitioners. They split the problem in two. Post-RS traceability follows a written requirement forward into design, code, and tests. Pre-RS traceability follows it backward to where it came from, who wanted it, and why. Their finding was that pre-RS traceability is where projects fail, because the origin is usually a conversation nobody recorded.

The Clarity Harness goes at pre-RS traceability by making the conversation the substrate, which is more ambitious than the demo lets on. It inherits the same limit: it only works for requirements that originated in something the harness ingested. A requirement that came from a hallway conversation, a competitor's pricing page, or a regulator's letter has no transcript to anchor to, and the trace will either be empty or point at the wrong source. The "widen to the source" button suggests one quoted sentence is often too thin an anchor on its own.

The mechanism description needs correcting

The video says "if you're using cloud code, it's going to spin up sub agents" (the auto-transcript mangles Claude Code), and "the orchestrator calls the agent, the agent calls the skill". Two corrections against Anthropic's documentation.

Anthropic writes subagents as one word. A subagent is a Markdown file with YAML frontmatter in .claude/agents/ or ~/.claude/agents/, and it runs in its own context window with its own system prompt, tool access, and permissions. It does not see the main conversation's history, which matters here: an interview agent and a requirement synthesizer running as subagents cannot silently share state, so the harness has to carry that state itself.

Skills do not work like function calls. An Agent Skill is a directory containing a SKILL.md file whose frontmatter name and description sit in the system prompt costing roughly 100 tokens, and Claude reads the body only when a request matches the description. Anthropic calls this progressive disclosure. So an agent does not "call a skill" the way it calls a tool. It matches a description and loads instructions. Claude Code's subagent frontmatter does support a skills: field that preloads specific skills, which is closer to what the video describes, but that is configuration rather than a runtime call. The distinction matters for anyone building on this: skill selection is a model decision, and model decisions need their own trace.

The judge is the bottleneck, and the screen admits it

The mission control pane in this demo shows 44 verdicts waiting against a session of 56 runs. Across the three visible agent groups, 58 runs have produced 36 pending verdicts with 12 reviewed. Those numbers do not quite reconcile with the header, which is normal for demo data, but the shape is the point. Runs accumulate faster than a human clears them, in a recording made by the person who built the tool.

"Then I'm the judge" is the load-bearing claim of the whole design, and it is the part with the least evidence behind it. Human judgment is also not a clean gold standard. Zheng and colleagues found in 2023 that GPT-4 agreed with expert human evaluators on about 85% of non-tie MT-Bench comparisons, close to the 81% agreement the human evaluators reached with each other. Putting a person in the loop buys accountability for the decision. It does not buy correctness, and it caps throughput at one person's attention span.

The harness has partial answers visible on screen. The auto_approved_dogfood label with a 7-day revert window and audit sampling is a real escape valve for volume. The Delta tab suggests version-to-version diffing exists. Neither gets a word of narration.

Saved feedback needs versioning or it drifts silently

The annotate box says your words become the training signal, and the narration says the feedback is saved so the next run is better. Agents carry a v0.1.0 label, so some version concept exists. Nothing visible ties a specific annotation to a specific version bump, and nothing shows what changed in the prompt or policy as a result.

That gap is where this class of system rots. Feedback accumulates, behavior shifts, and when output quality drops three months later there is no way to attribute the regression to the annotation that caused it. A feedback store that mutates agent behavior is a dependency, and dependencies need pinned versions, diffs, and rollback. The 7-day revert window covers artifacts. It does not appear to cover the accumulated judgment itself.

Key Takeaways

  • The harness is real internally and unavailable externally. heyclarity.dev reads "1,437 builders on the waitlist · launching soon · free & open source" and states that waitlist members "get the repo before it's public". Do not plan around it shipping.
  • Check github.com/Epistemic-Me before believing any open-source claim. The org has SDKs and scaffold repos, and no harness repo as of 2026-09-29.
  • If you want provenance on requirements today, the cheap version is a discipline, not a product: store the transcript, quote the sentence verbatim in the ticket, and link it. That covers most of what the Verdict tab does.
  • Build the review queue before the agent fleet. This demo shows 44 pending verdicts in a session with zero runs that day, which is the failure mode arriving early.
  • Version your feedback store the way you version prompts. Tag each annotation with the agent version it applied to, and keep a diff of what changed, or you cannot debug a regression later.
  • Use Anthropic's terms correctly when you design on top of them: subagents get isolated context windows, and skills load by description match rather than by explicit call.
  • Treat a 435-view founder demo as a design proposal. There is no independent user, no before-and-after measurement, and no third party in the recording.

Resources

Published September 28, 2026. Writeup generated from a favorited TikTok.