# Evidence-linked scoring | Call1 Technical Features

> How Call1 scores each rubric question on its own tier, and what every verdict carries: the evidence, the rationale, the confidence, and the versions in effect.

Source: https://call1.cc/features/evidence-linked-scoring

---

[CALL 1](https://call1.cc/)

[All features](https://call1.cc/features) [RFP reference](https://call1.cc/features#technical-reference) [Pricing](https://call1.cc/contact)

[Talk to us](https://call1.cc/contact)

[Features](https://call1.cc/features) / Evidence-linked scoring

03 / Rubrics, model routing, and auditability

# Answer any question about a score with the moment in the call that produced it.

How Call1 scores each rubric question on its own tier, and what every verdict carries with it: the evidence, the rationale, the confidence, and the versions in effect.

In this article

Why it matters

How it works

Technical shape

Data boundary

Performance

What it doesn't do

What to validate

## Why it matters

When an agent asks why a check failed, or an auditor asks how a rule was applied, you open the call. The verdict carries the transcript moment behind it, the rationale for the call it made, the confidence it had, and the rubric and model versions in effect when it ran. Two months and one model upgrade later, that record still reads the same way.

Your rubric holds different kinds of questions. A required disclosure is a phrase-matching problem. Empathy, objection handling, and save attempts need a model that works on meaning. A few questions, like whether a resolution fit the diagnosed problem or whether servicing language crossed into unfair or deceptive, are legal or professional judgment, and the starter packs ship those with human review attached. Each question is scored on its own tier, so paying large-model cost for a fixed disclosure phrase never happens.

One question at a time is also how you accept the system. Because each check is scored, versioned, and measured on its own, you can validate it against calls your team has already graded, and put five compliance checks into production before the rest of the rubric has cleared.

## How it works

1. **Publish a versioned rubric**
  Questions, sections, weights, applicability conditions, matchers, evidence requirements, and auto-fail behavior are published together as a version. Calls scored under that version keep it. A later edit changes what happens next, and leaves the existing results alone.
2. **Let the matchers set the tier**
  The matchers on a question determine which tier answers it: deterministic logic, a per-question adapter, or the local judgment model. An author sees the cost of a rule while writing it, and no one can move a question to an expensive tier without changing what it tests.
3. **Score each question**
  Every question returns pass, fail, flagged, or not applicable, with a confidence, a rationale, the transcript offsets that support it, the tier it used, and the base model and adapter versions behind it.
4. **Roll up the scorecard**
  Section weights and auto-fail behavior turn the question results into the scorecard your team reads. The individual results stay underneath the total, so any number on the scorecard decomposes into the checks that made it. A question marked not applicable leaves both sides of its section's arithmetic.
5. **Append the human decision**
  A reviewer confirms or overrides the verdict with a reason code. The machine result stays as it was written, and the adjudication lands beside it as a separate record. Reporting, exports, and the audit log keep the two distinguishable.

## Technical shape

### Deterministic tier

Exact and fuzzy phrase checks, required sequences, proximity, timing windows, counts, and metadata conditions. The matched span, or the place in the call where it should have been, is the entire explanation. Most disclosure and hold-time questions land here and need no training data.

### Adapter tier

Small per-question adapters served over one shared base model handle the semantic checks: acknowledging frustration, addressing an objection, attempting the save. Change your script and you retrain that question's adapter; the other forty stay where they are.

### Judgment tier

The local judgment model takes the smallest slice of the rubric, the questions that need broader context to answer at all. These arrive in a reviewer's queue for a human decision whatever confidence the model reports.

### Result and version contract

Question ID, verdict, confidence, rationale, evidence span, tier, rubric version, base model and adapter version, scoring time, and any adjudications that followed. That is the record an audit export carries, and nothing needed to read a two-year-old result is left implicit.

## Data boundary and privacy

> Every verdict resolves to evidence, rationale, the rubric configuration it ran against, and the model versions that produced it.

Scoring is the step where a model processes the words of the call. All three tiers run on your hardware inside your boundary. No transcript goes to a Call1-controlled inference endpoint.

An export can carry verdicts, metadata, and approved evidence excerpts with no audio attached. You decide whether transcript text may leave the deployment at all, and the answer can be no.

Machine verdicts are immutable and adjudications are append-only, so the file keeps what your reviewers saw at the time they decided. Rescoring under a newer rubric or adapter writes a new result next to the old one, and the original verdict, its evidence, and its versions stay readable.

## Performance

### Cost follows question difficulty

Phrase, timing, and metadata questions run as deterministic logic with no generative inference. Adapters carry the repeatable semantic checks over one resident base model. That leaves the judgment tier with the smallest workload on the box, so scoring every call fits on hardware you can buy once.

### Questions run independently

One call can mix all three tiers and run its checks concurrently, so a slow judgment question delays its own result and not the scorecard around it. Expedited calls come back in minutes; everything else is scored by morning.

### Immutable results cache well

Transcript, waveform peaks, evidence spans, and published verdicts do not change after processing, so the workbench caches them and fetches only the adjudication layer as it grows. A reviewer reopening a case gets it back in the state they left it.

## What it doesn't do

- No score reaches a screen without its rationale and a reference to the evidence it came from.
- Tier is read off the matchers on a question. It is not a cost or accuracy dial an author can turn.
- A reviewer's disagreement is recorded beside the machine verdict and never edits it.
- Per-question accuracy is validated against a sample of your own labeled calls before that check counts toward anything consequential for an agent.
- Scoring happens after the call. Call1 does not score, prompt, or intervene while a call is live.

## What to validate before deployment

1. Which of your rubric questions are phrase-matchable, which need an adapter, and which are judgment calls?
2. For each question, what must a passing result and a failing result show a reviewer?
3. Which labeled calls will you validate against, how many per question, and what precision, recall, and confidence calibration do you accept?
4. Which questions fail the whole call, which fail only their section, and which always route to a human?
5. Who may publish a rubric version, and what happens to reporting when one is published, rescored, or rolled back?

[Previous feature **Hybrid sentiment**](https://call1.cc/features/hybrid-sentiment) [Next feature **Smart review queue**](https://call1.cc/features/smart-review-queue)

## Evaluate it in your environment.

We'll map this feature to your data boundary, call volume, acceptance criteria, and operating responsibilities.

[Talk to us](https://call1.cc/contact)

[CALL 1](https://call1.cc/)

Every call. Inside your boundary.

[Pricing](https://call1.cc/contact)

[hello@call1.cc](mailto:hello@call1.cc)
