Jev vs LLMs for AI agent trace analysis: decisions, cost and evidence
Where Jev fits alongside LLMs in trace analysis, how to interpret the published comparison, and how to validate both on an original procurement example.

What is AI agent trace analysis?
AI agent trace analysis checks recorded model calls, tool actions and state changes to explain outcomes. This guide compares Jev's typed decisions with LLM-based interpretation, showing how to validate their labels, control cost and keep findings tied to evidence before a release.
Begin with the decision you need to make. For an agent that changes purchase orders, the question might be whether it respects approval and leaves the intended record state. For an operations sponsor, the useful output is a reviewed issue with affected work, evidence and an owner. A collection of interesting transcripts does not yet answer that question.
A trace is an instrumented record of observable execution. It can be incomplete and does not expose a model's private internal reasoning. OpenTelemetry describes spans as operations that can be related within traces.
The method below is educational guidance. Its event fields, classifier design, diagrams and calculations are proposed examples, not specifications of released Bibha features or reports of a customer deployment. The linked Bibha agent and monitoring pages provide platform context for the review method.
What is Jev, and what does it return?
Jev is TypeSafe AI's System One model. TypeSafe describes training with Reinforcement Learning for Calibrated Decisions (RLCD) and parallel typed outputs rather than generated prose. This interface constrains the answer shape; it does not make every semantic decision correct.
Its Choice questions select declared options, Score questions use defined levels, and Noul questions return a truth probability. Each question evaluates the same state independently. Choice and Score include confidence derived from their probability distributions. Noul has no separate confidence field.
Current documentation lists Jev 1.13, model ID jev-1.13.0, with text-only input, a 64k total request budget and a separate 32k limit for state plus the longest question. Pin the version and test the actual evidence window. Evaluate German and other non-English workloads separately.
Sources: TypeSafe AI: introducing System One models and Jev (external site)TypeSafe AI: Choice, Score and Noul questions (external site)TypeSafe AI: confidence and probability distributions (external site)TypeSafe AI: Jev models, versions and limits (external site)
Jev vs LLMs: choose the role before the model
Consider Jev for defined decisions once the policy, evidence view and labels are explicit. Consider an LLM for proposing new categories, interpreting unfamiliar cases or drafting an evidence-linked explanation. These are suggested roles to evaluate, not an assertion that one model wins every task.
LLMs can also produce schema-constrained structured values, as TypeSafe's introduction acknowledges. Compare the complete review process: accepted outputs, evidence quality, errors, abstentions, latency and actual cost. Keep human adjudication for important disagreements, and attach source-event references through validated application logic.
| Decision | Jev | LLM-based approach |
|---|---|---|
| Output | Declared typed decisions and distributions | Generated text or structured values |
| Suggested role | Apply defined review questions | Explore patterns and explain proposed findings |
| Validation | Check labels, probabilities and evidence links | Check labels, generated claims and evidence links |
| Context | Observe both versioned request limits | Check the selected model and request limits |
| Human review | Adjudicate uncertain or consequential decisions | Adjudicate uncertain or consequential decisions |
Sources: TypeSafe AI: introducing System One models and Jev (external site)TypeSafe AI: Jev models, versions and limits (external site)TypeSafe AI: Choice, Score and Noul questions (external site)TypeSafe AI: confidence and probability distributions (external site)
Read the published comparison in context
Applied Compute reported its 23 September 2026 study on 148 banking traces across 69 scenarios. Jev 1.13.0 at threshold 0.20 reached 85% recall against the union of Sol/Claude Opus labels, not human truth. Costs estimate one annotation over 10,000 traces. These experiment results are not Bibha results or a universal ranking; validate your workload.
| Model | Estimated cost |
|---|---|
| Jev | $11 |
| Luna | $54 |
| Haiku 4.5 | $479 |
Sources: Billion-Token Scale Trace Analysis: Jev vs LLMs | Applied Compute (external site)
Define the workload and the decision window
Measure scale in several units before choosing an analysis tool. Count eligible business runs, recorded events, retained bytes, analysis requests and input/output tokens separately. Record the time window, workflow mix and system versions beside those measures. Each answers a different capacity or quality question; one large token total cannot stand in for all of them.
Choose the unit that matches the business decision. A purchase-order task may cross several services and trace IDs, while one conversation may contain several orders. Link technical traces to an opaque task ID and define whether retries belong to the same run or a new attempt. Keep production tasks separate from generated test runs and training rollouts unless a comparison explicitly requires both.
For this guide, assume one eligible unit is one completed or interrupted order-change attempt. Its outcome is accepted only when the required approval and authoritative order state agree. This choice prevents a fast response or a successful HTTP request from becoming the success measure by accident. Other workflows need their own acceptance contract.
Write a short workload manifest before analysis: dates, inclusion rules, excluded runs, versions, retention gaps and the accountable owner. Track the distribution of trace lengths as well as the average. A small number of very long traces can determine request limits and review effort, so test those cases directly rather than hiding them inside a mean.
Record enough evidence to support a finding
An evidence contract defines what must be recorded to answer the review question. Start with stable task, run and event IDs; causal links; operation and result; workflow, model, tool and policy versions; and a reference to the outcome check. Add a completeness status so missing evidence is visible before classification begins.
OpenTelemetry provides timestamps, attributes, events, links and status around spans. W3C Trace Context specifies correlation headers, while current OpenTelemetry GenAI conventions remain labelled Development. Pin the conventions and instrumentation versions you use. The procurement fields proposed here are application fields, not a claim that a universal standard already defines every business approval record.
Specify the meaning of each field. An approval result should identify the action and record version it covers, its decision and the issuing authority. A mutation result should distinguish accepted, committed, rejected and uncertain outcomes. If a tool times out, preserve that ambiguity until a permitted state check resolves it. An absent span is not evidence that the operation never happened.
Test the collector contract with duplicate events, missing parents, delayed delivery and interrupted runs. Deduplicate by the documented event identity without discarding a legitimate second attempt. Store collector errors alongside completeness measures and compare accepted events with expected event families. Instrumentation health needs an owner because unreliable capture can look like unreliable agent behaviour.
| Field group | Example | Purpose |
|---|---|---|
| Identity | run_id = PO-DEMO-042; event_id = e04 | Connect a finding to a specific recorded action. |
| Causality | e04 follows e03; action = supplier_change | Check the prerequisite for the same action. |
| Versions | policy = P-7; workflow = W-2 | Reproduce the review under its original rules. |
| Outcome | commit = true; record_version = 18 | Separate tool response from business state. |
| Completeness | approval_event = present; payload = redacted | Show what the reviewer can and cannot establish. |
Sources: Traces | OpenTelemetry (external site)Trace Context | W3C Recommendation (external site)Semantic conventions for generative AI systems | OpenTelemetry (external site)
Query metadata before opening restricted content
Use a compact index to find relevant runs, then load only the permitted evidence needed for review. Keep identifiers, operation types, versions, timings and completeness searchable. Detailed prompts, retrieved documents and tool payloads may need a separate restricted store with its own retention and access rules. Their presence is a deliberate capture decision, not a default requirement.
The OpenTelemetry GenAI guidance treats input/output content capture as sensitive and opt-in. As a storage option, Parquet provides compressed columnar data and DuckDB supports projection and filter pushdown. Its partition handling can skip files selected out by partition filters. These are useful capabilities to assess in an existing data stack; they do not specify Bibha's implementation.
Start with the storage and query tools your team already operates. Partition by an appropriate time or workload boundary only when representative queries benefit. Keep raw personal or tenant identifiers out of file paths. A partition is a performance mechanism, so enforce organisation boundaries through actual access controls and test them independently.
Measure query latency, scan volume, ingestion delay and authorised content retrieval using a realistic window. Include deletion and expiry in the design: a source body may cease to be available while an annotation remains. Preserve the finding's permitted evidence references and record that availability change. Do not silently present a label as fully reproducible after its decisive source has expired.
Sources: Semantic conventions for generative client AI spans | OpenTelemetry (external site)Reading and Writing Parquet Files | DuckDB (external site)Hive Partitioning | DuckDB (external site)
A worked example: a supplier change before approval
The synthetic run PO-DEMO-042 illustrates an observable approval bypass. Policy P-7 requires a current approval before changing a purchase order's supplier. Event e03 records denial for that exact action; e04 records a committed change; e05 independently confirms the changed order. The approval and mutation refer to the same proposed change, so their relationship supports the finding.
The event table is deliberately small. All identifiers, policies, record versions and outcomes are invented. In a real investigation, verify the policy's applicability, approval validity, action identity and causal ordering before reaching the same conclusion. Wall-clock timestamps from separate machines may be imperfect; parent links and recorded sequencing can help establish what preceded the write.
Event e06 adds a separate problem: the final response describes the change as approved even though the recorded approval is denied. That is evidence about the response's claim. It does not explain why the system acted. A stale policy, orchestration defect or tool permission gap might be a cause, but each remains a hypothesis until the relevant component is tested.
If e03 were missing, the evidence would support a confirmed mutation but would not establish whether approval was granted. If e04 timed out and e05 were absent, the write outcome would remain uncertain. Those variants deserve different findings. The reviewer should not fill either gap with the confident wording of the final response.
| Event | Observed record | Interpretation |
|---|---|---|
| e01 | Request supplier_change for order version 17 | Defines the requested action. |
| e02 | P-7 requires current approval for supplier_change | Defines the prerequisite. |
| e03 | Approval denied for the requested change | The prerequisite is not satisfied. |
| e04 | supplier_change committed; record version 18 | A side effect occurred despite the denial. |
| e05 | Read confirms the supplier at version 18 | Confirms the recorded business state. |
| e06 | Response says the update was approved | The response conflicts with e03. |
Move through an evidence pipeline with review gates
A useful analysis pipeline separates capture, checking, interpretation and release authority. The diagram follows the synthetic purchase-order example through a proposed sequence: select the review window, check completeness, prepare permitted evidence, apply defined checks, adjudicate uncertain findings, and evaluate a proposed change. Each stage produces an inspectable output rather than an implied guarantee.
At the completeness gate, PO-DEMO-042 keeps its approval, mutation and state-check references. At the checking gate, the reviewer receives approval_bypass as a proposed label with e03 and e04 attached. At adjudication, the reviewer confirms the policy applies to this action. A release gate is reached only after an owned change has been tested on separate tasks under agreed acceptance criteria.
Keep the evidence-missing path visible. A run without the decisive approval record can enter a repair or human-review queue instead of being forced into success or failure. A run with a known transport error may pass a deterministic check but still need outcome verification. The visual branches represent these decisions; their size and position do not represent dataset proportions.
Log the analysis version and each gate decision. If a later policy or taxonomy revision changes an interpretation, retain the previous review record and explicitly create a new one. The pipeline then provides a way to reproduce decisions across a collection. It is a recommended operating pattern, with no claim that a diagram alone creates an automated or complete control system.
From recorded events to an evaluated change
Follow the evidence through six review stages. A proposed finding is checked before it becomes a change decision.
Stage 1
Review window
- Input
- Recorded events and workflow versions
Select the task family, time window and review question.
- Output
- A declared set of runs
Stage 2
Completeness
- Input
- Runs and required event families
Check event identity, causal links and decisive records.
- Output
- Evidence status and gap references
Missing decisive record? Repair capture or send for human review.
Stage 3
Permitted evidence
- Input
- Authorised records and access rules
Keep decisive references; mark redactions and omissions.
- Output
- A compact, linked evidence view
Stage 4
Defined checks
- Input
- Evidence and versioned label definitions
Apply a suitable rule or classifier; cite supporting events.
- Output
- Proposed labels with references
Unsupported interpretation? Abstain and request review.
Stage 5
Adjudication
- Input
- Proposed labels, policy and event references
Resolve exceptions and disagreements with an accountable reviewer.
- Output
- A reviewed finding or an unresolved case
Stage 6
Change evaluation
- Input
- An owned candidate change and separate test tasks
Compare acceptance results with the baseline and record limits.
- Output
- Evidence for a release decision
Acceptance unmet? Revise the change before seeking release approval.
Reduce input without removing the decisive event
Preprocessing should make relevant evidence easier to inspect while preserving the basis of the finding. Build a compact operational view from stable event references, permitted excerpts, action identity, prerequisite results and side-effect outcomes. Keep the source record within its authorised retention boundary. Mark every redaction, omission and truncation so the reader can judge the view's limits.
For PO-DEMO-042, the compact view must retain e03's denied approval and e04's committed change. Removing a repeated catalogue response may be harmless to this question. Removing the approval because it looks like routine metadata changes the question itself. A summary saying that the order was updated would erase the evidence needed to distinguish an approved update from a bypass.
Lost in the Middle found that information position affected the tested retrieval and question-answering tasks. LongLLMLingua studies prompt compression in specified long-context tasks. These papers motivate testing how evidence selection affects your classifier; they do not establish that a particular current model will miss an event or that compression preserves every relevant detail.
Create a review set with long traces, late prerequisites, repeated tool outputs and contradictory messages. Compare labels and supporting references before and after reduction. Inspect disagreements and record whether the missing evidence was relevant. Version the reduction procedure with the rubric. A lower token bill is useful only when the resulting view still supports the decisions the team intends to make.
Sources: Lost in the Middle: How Language Models Use Long Contexts (external site)LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression (external site)
Create failure labels with boundaries and owners
A failure taxonomy is a versioned vocabulary for observed patterns. Start by reviewing varied runs with the people who own the workflow, then draft a small set of labels that answer operational questions. Give each label inclusion evidence, exclusions, severity and an owner. Keep observations, proposed causes and business consequences in separate fields.
The interactive taxonomy uses an original six-label procurement example. Selecting a label should reveal its meaning and the evidence needed to apply it. The table provides the same definitions in readable text. These categories are teaching examples, not a standard, a customer result or a released Bibha classifier. A run may carry several labels, while unknown means the current definitions cannot settle its interpretation.
For approval_bypass, require evidence that the applicable approval was denied or otherwise unsatisfied when the specific change committed. A missing approval event alone is excluded because the collection may be incomplete. evidence_missing records that limitation. This boundary stops a data-capture defect from being automatically reported as a policy violation and gives the instrumentation owner a separate issue to resolve.
Use discovery reviews to improve definitions before measuring a release. Record the policy and taxonomy version, keep difficult examples, and freeze the vocabulary for each comparison window. If a category later splits, document the mapping and re-score a common comparison set where needed. New labels can be useful, but their arrival should not be mistaken for a deterioration in the agent.
| Label | Include when | Exclude when | Illustrative severity | Review owner |
|---|---|---|---|---|
| approval_bypass | A change commits despite an applicable unsatisfied approval. | Approval evidence is absent or belongs to a different action. | High for this denied supplier change | Workflow and approval owner |
| unverified_state | A completion claim lacks the state evidence required by the acceptance contract. | The required independent state check is present and agrees. | Depends on the claimed outcome | Integration owner |
| duplicate_action | Two committed side effects exceed the permitted single action. | A retry is rejected or the operation is safely idempotent. | High if the extra side effect is consequential | Integration owner |
| evidence_missing | Required records are missing, truncated or unavailable. | All evidence needed for this finding is available. | Evidence limitation; operational risk unresolved | Instrumentation owner |
| tool_error | A recorded tool operation returns an error or uncertain transport outcome. | Only the business outcome is wrong with no recorded tool error. | Depends on recovery and final state | Tool owner |
| unknown | Available evidence cannot be assigned reliably under this taxonomy. | A defined label is supported or required evidence is missing. | Review required before assigning risk | Domain reviewer |
Which finding does the evidence support?
Expand an example to compare its definition, boundary and observed events. A run can support several labels; these cases isolate different decisions.
Synthetic teaching examples, not customer data or a benchmark
approval_bypassA denied change still commits
- Definition
- A change commits despite an applicable approval requirement being unsatisfied.
- Include when
- The denial and committed change refer to the same action under the applicable policy, with no valid exception.
- Exclude when
- Approval evidence is missing, applies to another action, or a valid policy exception permits the change.
Synthetic trace: PO-DEMO-042
e03P-7 approval denied for the requested supplier change.e04That supplier change commits at order version 18.e05Independent read-back verifies the changed supplier at version 18.
- Supported finding
- approval_bypass only. The read-back verifies state, so this example does not support unverified_state.
- What to review
- Confirm policy P-7, action identity and causal order; inspect where the workflow should enforce the approval requirement.
unverified_stateCompletion is claimed before state is checked
- Definition
- A completion claim lacks the state evidence required by the acceptance contract.
- Include when
- A fully captured workflow claims completion but does not perform the required independent state check.
- Exclude when
- The required state check is present and agrees with the completion claim. An accepted request alone is not a state check.
Synthetic trace: UPDATE-DEMO-017
e01Write request accepted for asynchronous processing.e02Workflow ends without invoking the required read-back; capture is complete.e03Final response says the record was updated.
- Supported finding
- unverified_state. The final business state remains unresolved. The verification step was omitted, rather than lost by the collector.
- What to review
- Confirm the acceptance contract and run completeness; obtain a permitted state check before treating the update as complete.
duplicate_actionA retry creates a second side effect
- Definition
- Two committed side effects exceed the permitted single action.
- Include when
- Distinct operation results confirm two committed effects for one action that permits only one.
- Exclude when
- The retry is rejected, reuses an idempotent result, or the collector records the same event twice.
Synthetic trace: RESERVE-DEMO-016
e01Workflow authorises one reservation.e02First operation commits reservation R-DEMO-201.e03Retry commits a distinct reservation R-DEMO-202.
- Supported finding
- duplicate_action. Two distinct committed results support the finding; the number of attempts alone would not.
- What to review
- Confirm the single-action constraint, operation identities and resulting records; inspect idempotency and retry handling.
evidence_missingThe decisive approval record is unavailable
- Definition
- Required evidence is missing, truncated or unavailable for the review question.
- Include when
- The collector or permitted evidence view cannot supply the record needed to establish the approval decision.
- Exclude when
- All decisive records are available. An adverse outcome does not by itself establish a capture gap.
Synthetic trace: CAPTURE-DEMO-023
e01Collector marks the required approval record as lost.e02Supplier change is recorded as committed.e03Read-back confirms the changed supplier.
- Supported finding
- evidence_missing. The mutation is verified, but the trace cannot establish whether approval was granted or denied.
- What to review
- Recover authorised approval evidence and inspect collector loss; keep approval_bypass undecided until its prerequisite evidence exists.
tool_errorA timeout leaves the write outcome uncertain
- Definition
- A tool operation records an error or an uncertain transport outcome.
- Include when
- The tool records a definite failure, or a transport timeout whose business outcome has not yet been resolved.
- Exclude when
- Only the business outcome is wrong and no tool error is recorded. A verified later success does not erase a recorded transport error.
Synthetic trace: WRITE-DEMO-031
e01Write request is sent.e02Transport times out without a commit result.e03Run closes with write outcome explicitly unresolved; capture is complete.
- Supported finding
- tool_error with an uncertain outcome. This timeout does not prove rejection or commit; a definite validation rejection would prove that attempt failed.
- What to review
- Inspect operation identity and retry safety, then perform a permitted state check. Do not retry blindly when the first write may have committed.
unknownComplete evidence reaches an undefined exception
- Definition
- Available evidence cannot be assigned reliably under the current taxonomy.
- Include when
- The decisive records are available, but an unfamiliar policy interpretation is outside the defined label boundaries.
- Exclude when
- A defined label is supported, or missing decisive records prevent interpretation. Missing records belong under evidence_missing.
Synthetic trace: EXCEPTION-DEMO-008
e01Policy returns delegated_exception with its complete decision record.e02The scoped change commits and read-back confirms the resulting state.e03Collector confirms that decision, write and state-check records are complete.
- Supported finding
- unknown pending domain review: the rubric does not define delegated_exception. The gap is in interpretation, not in the recorded evidence; bypass is not established.
- What to review
- Ask the policy owner to resolve the exception, document the decision and version the taxonomy before comparing later results.
Match the check to the available evidence
Use deterministic checks when the relevant condition is explicit in structured data. Use semantic classification when interpretation depends on policy language or several related events. Use human adjudication for disputed, unfamiliar or consequential cases. These methods can work together; none should inherit execution authority simply because it produces a label.
A deterministic rule can compare e03's decision with e04's commit state if their action IDs and policy scope are reliable. An LLM-based semantic classifier can propose that e06 contradicts the recorded approval, returning the supporting event IDs and a short explanation. A domain reviewer can resolve whether an exceptional approval path applies. Require the same evidence discipline from all three.
Model-based judges need their own checks. The MT-Bench judge study identifies position, verbosity and self-enhancement biases in its evaluated settings. For this workflow, test whether reordering equivalent evidence, changing response length or naming the model alters labels. Those experiments establish limits for the chosen judge and rubric instead of borrowing another benchmark's reported agreement.
Keep classifier input and output constrained. Supply the permitted evidence, applicable policy and current definitions; validate returned label IDs and event references. An explanation without a valid reference is a review suggestion, not an established finding. When evidence is unavailable or interpretation conflicts, let the checker abstain and route the case to the appropriate queue.
| Method | Useful for | Limit | Required output |
|---|---|---|---|
| Deterministic | Explicit denial plus matching committed action; duplicate IDs; missing fields | Only as sound as the recorded fields and rule scope | Rule version, result and event references |
| Semantic | Policy interpretation, contradictions and cross-event meaning | Can misread evidence or produce unsupported explanations | Rubric version, proposed labels, references and abstention |
| Human | Exceptions, disagreement, new patterns and consequential findings | Review capacity and consistent guidance are required | Decision, rationale, evidence and accountable reviewer |
Sources: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (external site)
Build reviewed examples before measuring a classifier
Create a reference set that represents the decisions you want the classifier to make. Include accepted runs, relevant failures, incomplete evidence and legitimate exceptions. Ask domain reviewers to label the evidence independently where practical, then adjudicate disagreements. Their resolved decisions become the reference for that dataset, with limits recorded rather than described as infallible truth.
Separate label discovery, development, threshold tuning and the final test. Keep related attempts, duplicate tasks and events from one run in the same split. Where release generalisation matters, reserve a later time window or a task family that development has not touched. The split should reflect the future decision, not merely make the spreadsheet convenient.
For procurement, include approval granted, approval denied, approval expired, a permitted exception, a timed-out write with confirmed state and a timed-out write with unresolved state. Add a duplicate request whose idempotency key prevents a second mutation. These cases check whether the definitions distinguish superficially similar events with different operational meanings.
Hide reference diagnoses and expected labels from the classifier's inputs. Keep a record of how examples were selected and which cases lack a settled reference. If the team changes a label definition after reading final-test mistakes, that set has become development material. Reserve fresh evaluation evidence before claiming the revised system generalises.
Report precision, recall and calibrated decisions separately
Measure each important label, including its support and review burden. Precision asks how many proposed positives are correct; recall asks how many reference positives were found. For an explicitly synthetic result with 80 true positives, 20 false positives and 40 false negatives, precision is 80% and recall is about 66.7%. These are arithmetic examples, not measured classifier results. Apply these checks to both Jev decisions and LLM-generated labels.
Report the confusion counts, dataset window and per-label sample size beside those ratios. A rare approval bypass can disappear inside overall accuracy when most runs are acceptable. Multi-label evaluation also needs a stated counting unit. If no relevant positives or predictions exist, report the score as undefined where appropriate and explain the available evidence instead of inventing a perfect value.
Calibration is a separate property. It compares predicted probabilities with observed label frequencies; a model's written confidence score is not automatically a probability. Fit or assess calibration using appropriate separate data, and tune action thresholds away from the untouched final test. The scikit-learn documentation distinguishes probability estimation, calibration and threshold choice.
Choose thresholds according to the action. An investigation queue may accept more false alarms to find additional cases, while a consequential operational decision requires stricter evidence and review. Show what each threshold does to missed cases and reviewer capacity. Recheck performance when the model, rubric, policy, preprocessing or workload changes; the previous measurement describes its recorded configuration.
Sources: Precision, recall, F-score and support | scikit-learn (external site)Probability calibration | scikit-learn (external site)Tuning the decision threshold for class prediction | scikit-learn (external site)
Sample for coverage and keep uncertainty visible
Separate retention sampling from investigation sampling and population measurement. OpenTelemetry describes head sampling as an early decision and tail sampling as a decision using later trace evidence. Its tail-sampler implementation has routing, memory and late-span constraints. A configured error policy cannot establish that every failing run was retained, especially if evidence was dropped upstream.
Maintain a probability-sampled baseline where feasible, alongside targeted cohorts for rare events, new versions, slow runs and incomplete capture. Record why each run was selected and its effective inclusion probability when that probability is valid. OpenTelemetry's probability-sampling guidance defines adjusted counts under stated assumptions; targeted selections without known probabilities do not acquire valid weights merely because they are useful for debugging.
A collection deliberately enriched for approval errors is suitable for investigating those errors. Its raw label percentage is not the production failure rate. Population estimates need the correct denominator, valid selection weights, duplicate handling and consideration of classification errors. Report the target population and the uncertainty around the estimate rather than presenting a precise number without a defensible design.
Keep unknown and evidence_missing visible in reporting. Count their review queues, ageing and resolution separately from confirmed failures and accepted outcomes. Do not silently call abstention success. In multi-label reporting, label counts can exceed run counts, so show both. Rising unknowns may indicate a new workflow, a weak definition or changed evidence capture; review examples before selecting the remedy.

Sources: Sampling | OpenTelemetry (external site)Tail Sampling Processor | OpenTelemetry Collector Contrib (external site)TraceState: Probability Sampling | OpenTelemetry (external site)
Estimate token cost and the work around it
Estimate cost from the actual analysis plan, including selection and repeated passes. Let N be eligible retained traces, s the selection fraction and k the passes per selected trace. Multiply N × s × k by average input and output tokens per pass. Price non-overlapping billed categories using the same currency and analysis window.
Here is a wholly illustrative example: 100,000 traces × 0.25 selected × 2 passes equals 50,000 passes. At 1,500 input and 100 output tokens per pass, that is 75 million input tokens and 5 million output tokens. Assume 1 cost unit per million input tokens and 4 per million output tokens. Classification then costs 75 × 1 + 5 × 4 = 95 units.
These volumes and prices are assumptions, not Jev pricing, provider tariffs, Bibha results or a quote. The example excludes retries, caching, request rounding and differing trace lengths until those terms are added. Google publishes distinct billing categories for its models and modes; consult the actual provider contract when building an estimate. Avoid adding cached-input or reasoning-detail counters again when they are already included in billed totals.
Add discovery work, data preparation, storage, scanning, deterministic compute, orchestration and human review separately. Test sensitivity to longer evidence, more passes and a larger selected fraction. Then ask whether the reduced budget still covers important failures and uncertain cases. A cheap classifier that overwhelms reviewers or omits the decisive event can make the complete operating process more expensive.
| Term | Calculation | Result |
|---|---|---|
| Selected traces | 100,000 × 0.25 | 25,000 |
| Classifier passes | 25,000 × 2 | 50,000 |
| Input tokens | 50,000 × 1,500 | 75,000,000 |
| Output tokens | 50,000 × 100 | 5,000,000 |
| Input cost | 75 × 1 | 75 units |
| Output cost | 5 × 4 | 20 units |
| Classification total | 75 + 20 | 95 units before other costs |
C_classification = N * s * k * ((t_in / 1000000) * p_in + (t_out / 1000000) * p_out)
C_total = C_classification + C_discovery + C_data_and_storage + C_deterministic_compute + C_orchestration + C_human_reviewSources: Gemini Developer API pricing | Google AI for Developers (external site)
Treat trace content as untrusted evaluator input
Protect the analysis system as carefully as the agent being reviewed. Trace bodies can contain customer text, retrieved documents and tool responses that an attacker influences. AgentDojo evaluates prompt-injection attempts delivered through untrusted tool content. This supports testing the evaluator's trust boundary separately from ordinary task-quality checks; it does not establish an immune defence.
For the proposed procurement review, give the evaluator read-only access to the permitted evidence view and no authority to change orders. Treat embedded requests to ignore policy or reveal data as content to analyse. Keep evaluator instructions, policy and source evidence clearly separated. Validate the output against allowed labels and verify that cited events belong to the authorised run.
Minimise content capture and agree access, processing location, retention and deletion with the system owner. Redact secrets and personal details where required before export, with redaction marked in the evidence view. W3C Trace Context forbids personal or sensitive information in correlation headers. Trace IDs should connect records; they should not serve as permission to retrieve them.
Test fabricated event references, malicious retrieved text, cross-organisation IDs and attempts to trigger evaluator tools. Include incomplete and redacted traces in those tests. Record utility and boundary failures independently so an apparently helpful diagnosis cannot conceal an access violation. Keep logs free of unnecessary source bodies and provide a controlled route for authorised reviewers to inspect the underlying evidence.
Sources: AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents (external site)Trace Context | W3C Recommendation (external site)
Turn a reviewed finding into a controlled release
Use trace findings to justify an owned change and a repeatable acceptance test. For PO-DEMO-042, the first task is to confirm where the approval requirement should be enforced, then test the smallest relevant change. A proposed label does not authorise a production mutation. The workflow owner and release authority remain responsible for the decision.
Evaluate the complete task, including final state. tau-bench demonstrates checking a post-interaction database state against an annotated goal and assessing repeated success. Apply that idea carefully to your own acceptance contract: a denied approval should leave the supplier unchanged, an approved change should reach the intended state, and a retry should not create a duplicate side effect.
Compare the baseline and candidate on a fixed set, a held-out set and the required adversarial cases. Keep task mix, policy, taxonomy and evaluator versions constant where the comparison needs them. Report quality, latency, resource use and unresolved cases together. If the measurement procedure changes, separate that change from the system improvement you are trying to establish.
Before release, name the approver, rollout scope, alert owner and rollback trigger. Check recovery from a partial rollout or failed dependency. After release, inspect fresh runs, the sampled baseline and unresolved queues under the same declared contract. Record what changed, the evidence supporting it and what remains uncertain. A controlled release closes one review cycle while preserving the means to discover the next issue.
- A practical release record. Record the review window and versions, approved evidence access, adjudicated findings, change owner, acceptance results, remaining limitations, release authority and rollback plan. Attach the permitted evidence references so the next reviewer can assess the decision.
Sources: tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (external site)
Connect the evidence to the workflow you operate
Bibha helps teams connect task traces and operating measures with evaluation, reviewed feedback and controlled improvement releases. Its monitoring, evaluation and governance offering covers agreed business outcomes, task-specific tests, human review, versioned results and defined access and release responsibilities. The public platform pages explain that scope; implementation details depend on the system and engagement.
The event contract, storage choices, taxonomy and classifier pipeline in this guide are educational recommendations. They should be assessed against your workflow, evidence permissions and operating constraints before becoming a design. No Bibha throughput, classifier score, saving or automatic production-learning result is claimed here. Monitoring evidence helps people make a reviewed change; it does not mean every interaction changes the live model.
Prepare a workflow brief with the decision you want to improve, the systems involved, the authoritative outcome and the approval owner. Add the evidence you are permitted to retain, the review volume and the operating budget. Those inputs make it possible to discuss the actual boundary between platform configuration, integration work and ongoing review.
If you are planning an agent workflow, bring that brief to a conversation with Bibha. Start with one consequential question, such as whether a change can occur without current approval, and agree how its evidence will be checked. A useful first outcome is a scoped review and release plan with named responsibilities and acceptance criteria.
Questions and answers
Is Jev a Bibha or Applied Compute model?
Jev is a model from TypeSafe AI. This guide assesses its possible role in trace review without claiming Bibha owns Jev or provides an approved Jev integration.
Does Jev confidence prove that a trace label is correct?
No. For Choice and Score, confidence is derived from the returned probability distribution. Noul has no separate confidence field. Test the questions, evidence and version on reviewed examples and retain appropriate human adjudication.
How do traces differ from ordinary logs?
Logs record individual messages or events. A trace links recorded operations within an execution path. For business review, connect those technical records to the task and its authoritative outcome. Both can contribute evidence; neither is complete simply because it exists.
Can an LLM classify every failure reliably?
No reliability claim follows from using an LLM. Evaluate the chosen model, evidence view and rubric on reviewed, held-out examples. Report per-label errors and abstentions, and retain human review for uncertain or consequential findings.
What should happen when an approval event is missing?
Mark the evidence gap and check permitted source records or capture health. A missing approval event does not prove denial or approval. In the procurement example, reserve approval_bypass for a supported unsatisfied prerequisite and committed matching action.
Is a confidence score a calibrated probability?
Only if appropriate evaluation supports that interpretation. A generated confidence number alone is insufficient. Assess its relationship to observed label frequencies on separate data, and choose action thresholds according to the decision and its consequences.
Can a targeted sample show the production failure rate?
Its raw percentage cannot establish that rate. Targeted incident samples deliberately change the mix. Population estimates require a defined population, suitable inclusion probabilities or sampling design, defensible denominators and consideration of classifier error.
How can we lower analysis cost without losing evidence?
Query metadata first, select relevant permitted events, remove repetition carefully and measure the effect on findings. Count passes and retries, then add storage, compute and review costs. Preserve decisive evidence and mark every omission.
Does analysing traces reveal hidden model reasoning?
No. This guide uses instrumented operations, recorded content and observable state. Explanations about causes remain hypotheses until tested; a trace does not provide access to private internal reasoning.
Should labels automatically change a production agent?
A label is analysis output. Turn it into an owned issue, verify the cause, test a proposed change and obtain the accountable release decision. Keep release authority separate from the diagnostic evaluator.
References
- Billion-Token Scale Trace Analysis: Jev vs LLMs | Applied Compute https://www.appliedcompute.com/platform/billion-token-scale-trace-analysis (external site)
- Traces | OpenTelemetry https://opentelemetry.io/docs/concepts/signals/traces/ (external site)
- Trace Context | W3C Recommendation https://www.w3.org/TR/trace-context/ (external site)
- Semantic conventions for generative AI systems | OpenTelemetry https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/README.md (external site)
- Semantic conventions for generative client AI spans | OpenTelemetry https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md (external site)
- Sampling | OpenTelemetry https://opentelemetry.io/docs/concepts/sampling/ (external site)
- Tail Sampling Processor | OpenTelemetry Collector Contrib https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/processor/tailsamplingprocessor/README.md (external site)
- TraceState: Probability Sampling | OpenTelemetry https://opentelemetry.io/docs/specs/otel/trace/tracestate-probability-sampling/ (external site)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena https://arxiv.org/abs/2306.05685 (external site)
- Precision, recall, F-score and support | scikit-learn https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_recall_fscore_support.html (external site)
- Probability calibration | scikit-learn https://scikit-learn.org/stable/modules/calibration.html (external site)
- Tuning the decision threshold for class prediction | scikit-learn https://scikit-learn.org/stable/modules/classification_threshold.html (external site)
- Lost in the Middle: How Language Models Use Long Contexts https://arxiv.org/abs/2307.03172 (external site)
- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression https://arxiv.org/abs/2310.06839 (external site)
- tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains https://arxiv.org/abs/2406.12045 (external site)
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents https://arxiv.org/abs/2406.13352 (external site)
- Reading and Writing Parquet Files | DuckDB https://duckdb.org/docs/current/data/parquet/overview (external site)
- Hive Partitioning | DuckDB https://duckdb.org/docs/current/data/partitioning/hive_partitioning (external site)
- Gemini Developer API pricing | Google AI for Developers https://ai.google.dev/gemini-api/docs/pricing (external site)
- TypeSafe AI: introducing System One models and Jev https://typesafe.ai/blog/introducing-system-one-models-and-jev (external site)
- TypeSafe AI: Jev models, versions and limits https://docs.typesafe.ai/models (external site)
- TypeSafe AI: Choice, Score and Noul questions https://docs.typesafe.ai/primitives (external site)
- TypeSafe AI: confidence and probability distributions https://docs.typesafe.ai/confidence (external site)