Skip to main content
Bibha home

Engineering guide,

Jev vs LLMs for AI agent trace analysis: decisions, cost and evidence

Where Jev fits alongside LLMs in trace analysis, how to interpret the published comparison, and how to validate both on an original procurement example.

By Prashant Kumar, Founder & CEO Bibha Ai Labs

A small evidence sample expands towards a larger field of classification marks through a sparse green cone diagram.
Original AI-generated conceptual illustration of scaling a reviewed sample into a broader classification task. The geometry is explanatory, not measured data.

What is AI agent trace analysis?

AI agent trace analysis checks recorded model calls, tool actions and state changes to explain outcomes. This guide compares Jev's typed decisions with LLM-based interpretation, showing how to validate their labels, control cost and keep findings tied to evidence before a release.

Begin with the decision you need to make. For an agent that changes purchase orders, the question might be whether it respects approval and leaves the intended record state. For an operations sponsor, the useful output is a reviewed issue with affected work, evidence and an owner. A collection of interesting transcripts does not yet answer that question.

A trace is an instrumented record of observable execution. It can be incomplete and does not expose a model's private internal reasoning. OpenTelemetry describes spans as operations that can be related within traces.

The method below is educational guidance. Its event fields, classifier design, diagrams and calculations are proposed examples, not specifications of released Bibha features or reports of a customer deployment. The linked Bibha agent and monitoring pages provide platform context for the review method.

Sources: Traces | OpenTelemetry (external site)

What is Jev, and what does it return?

Jev is TypeSafe AI's System One model. TypeSafe describes training with Reinforcement Learning for Calibrated Decisions (RLCD) and parallel typed outputs rather than generated prose. This interface constrains the answer shape; it does not make every semantic decision correct.

Its Choice questions select declared options, Score questions use defined levels, and Noul questions return a truth probability. Each question evaluates the same state independently. Choice and Score include confidence derived from their probability distributions. Noul has no separate confidence field.

Current documentation lists Jev 1.13, model ID jev-1.13.0, with text-only input, a 64k total request budget and a separate 32k limit for state plus the longest question. Pin the version and test the actual evidence window. Evaluate German and other non-English workloads separately.

Sources: TypeSafe AI: introducing System One models and Jev (external site)TypeSafe AI: Choice, Score and Noul questions (external site)TypeSafe AI: confidence and probability distributions (external site)TypeSafe AI: Jev models, versions and limits (external site)

Jev vs LLMs: choose the role before the model

Consider Jev for defined decisions once the policy, evidence view and labels are explicit. Consider an LLM for proposing new categories, interpreting unfamiliar cases or drafting an evidence-linked explanation. These are suggested roles to evaluate, not an assertion that one model wins every task.

LLMs can also produce schema-constrained structured values, as TypeSafe's introduction acknowledges. Compare the complete review process: accepted outputs, evidence quality, errors, abstentions, latency and actual cost. Keep human adjudication for important disagreements, and attach source-event references through validated application logic.

Practical selection criteria. Model outputs still require workload-specific validation.
DecisionJevLLM-based approach
OutputDeclared typed decisions and distributionsGenerated text or structured values
Suggested roleApply defined review questionsExplore patterns and explain proposed findings
ValidationCheck labels, probabilities and evidence linksCheck labels, generated claims and evidence links
ContextObserve both versioned request limitsCheck the selected model and request limits
Human reviewAdjudicate uncertain or consequential decisionsAdjudicate uncertain or consequential decisions

Sources: TypeSafe AI: introducing System One models and Jev (external site)TypeSafe AI: Jev models, versions and limits (external site)TypeSafe AI: Choice, Score and Noul questions (external site)TypeSafe AI: confidence and probability distributions (external site)

Read the published comparison in context

Applied Compute reported its 23 September 2026 study on 148 banking traces across 69 scenarios. Jev 1.13.0 at threshold 0.20 reached 85% recall against the union of Sol/Claude Opus labels, not human truth. Costs estimate one annotation over 10,000 traces. These experiment results are not Bibha results or a universal ranking; validate your workload.

Study estimates, not current tariffs.
ModelEstimated cost
Jev$11
Luna$54
Haiku 4.5$479

Sources: Billion-Token Scale Trace Analysis: Jev vs LLMs | Applied Compute (external site)

Define the workload and the decision window

Measure scale in several units before choosing an analysis tool. Count eligible business runs, recorded events, retained bytes, analysis requests and input/output tokens separately. Record the time window, workflow mix and system versions beside those measures. Each answers a different capacity or quality question; one large token total cannot stand in for all of them.

Choose the unit that matches the business decision. A purchase-order task may cross several services and trace IDs, while one conversation may contain several orders. Link technical traces to an opaque task ID and define whether retries belong to the same run or a new attempt. Keep production tasks separate from generated test runs and training rollouts unless a comparison explicitly requires both.

For this guide, assume one eligible unit is one completed or interrupted order-change attempt. Its outcome is accepted only when the required approval and authoritative order state agree. This choice prevents a fast response or a successful HTTP request from becoming the success measure by accident. Other workflows need their own acceptance contract.

Write a short workload manifest before analysis: dates, inclusion rules, excluded runs, versions, retention gaps and the accountable owner. Track the distribution of trace lengths as well as the average. A small number of very long traces can determine request limits and review effort, so test those cases directly rather than hiding them inside a mean.

Record enough evidence to support a finding

An evidence contract defines what must be recorded to answer the review question. Start with stable task, run and event IDs; causal links; operation and result; workflow, model, tool and policy versions; and a reference to the outcome check. Add a completeness status so missing evidence is visible before classification begins.

OpenTelemetry provides timestamps, attributes, events, links and status around spans. W3C Trace Context specifies correlation headers, while current OpenTelemetry GenAI conventions remain labelled Development. Pin the conventions and instrumentation versions you use. The procurement fields proposed here are application fields, not a claim that a universal standard already defines every business approval record.

Specify the meaning of each field. An approval result should identify the action and record version it covers, its decision and the issuing authority. A mutation result should distinguish accepted, committed, rejected and uncertain outcomes. If a tool times out, preserve that ambiguity until a permitted state check resolves it. An absent span is not evidence that the operation never happened.

Test the collector contract with duplicate events, missing parents, delayed delivery and interrupted runs. Deduplicate by the documented event identity without discarding a legitimate second attempt. Store collector errors alongside completeness measures and compare accepted events with expected event families. Instrumentation health needs an owner because unreliable capture can look like unreliable agent behaviour.

Proposed application evidence fields for the synthetic procurement workflow. These are recommendations, not a standardised SDK schema.
Field groupExamplePurpose
Identityrun_id = PO-DEMO-042; event_id = e04Connect a finding to a specific recorded action.
Causalitye04 follows e03; action = supplier_changeCheck the prerequisite for the same action.
Versionspolicy = P-7; workflow = W-2Reproduce the review under its original rules.
Outcomecommit = true; record_version = 18Separate tool response from business state.
Completenessapproval_event = present; payload = redactedShow what the reviewer can and cannot establish.

Sources: Traces | OpenTelemetry (external site)Trace Context | W3C Recommendation (external site)Semantic conventions for generative AI systems | OpenTelemetry (external site)

Query metadata before opening restricted content

Use a compact index to find relevant runs, then load only the permitted evidence needed for review. Keep identifiers, operation types, versions, timings and completeness searchable. Detailed prompts, retrieved documents and tool payloads may need a separate restricted store with its own retention and access rules. Their presence is a deliberate capture decision, not a default requirement.

The OpenTelemetry GenAI guidance treats input/output content capture as sensitive and opt-in. As a storage option, Parquet provides compressed columnar data and DuckDB supports projection and filter pushdown. Its partition handling can skip files selected out by partition filters. These are useful capabilities to assess in an existing data stack; they do not specify Bibha's implementation.

Start with the storage and query tools your team already operates. Partition by an appropriate time or workload boundary only when representative queries benefit. Keep raw personal or tenant identifiers out of file paths. A partition is a performance mechanism, so enforce organisation boundaries through actual access controls and test them independently.

Measure query latency, scan volume, ingestion delay and authorised content retrieval using a realistic window. Include deletion and expiry in the design: a source body may cease to be available while an annotation remains. Preserve the finding's permitted evidence references and record that availability change. Do not silently present a label as fully reproducible after its decisive source has expired.

Sources: Semantic conventions for generative client AI spans | OpenTelemetry (external site)Reading and Writing Parquet Files | DuckDB (external site)Hive Partitioning | DuckDB (external site)

A worked example: a supplier change before approval

The synthetic run PO-DEMO-042 illustrates an observable approval bypass. Policy P-7 requires a current approval before changing a purchase order's supplier. Event e03 records denial for that exact action; e04 records a committed change; e05 independently confirms the changed order. The approval and mutation refer to the same proposed change, so their relationship supports the finding.

The event table is deliberately small. All identifiers, policies, record versions and outcomes are invented. In a real investigation, verify the policy's applicability, approval validity, action identity and causal ordering before reaching the same conclusion. Wall-clock timestamps from separate machines may be imperfect; parent links and recorded sequencing can help establish what preceded the write.

Event e06 adds a separate problem: the final response describes the change as approved even though the recorded approval is denied. That is evidence about the response's claim. It does not explain why the system acted. A stale policy, orchestration defect or tool permission gap might be a cause, but each remains a hypothesis until the relevant component is tested.

If e03 were missing, the evidence would support a confirmed mutation but would not establish whether approval was granted. If e04 timed out and e05 were absent, the write outcome would remain uncertain. Those variants deserve different findings. The reviewer should not fill either gap with the confident wording of the final response.

Invented events for PO-DEMO-042, shown in causal order. No customer data or production measurements are used.
EventObserved recordInterpretation
e01Request supplier_change for order version 17Defines the requested action.
e02P-7 requires current approval for supplier_changeDefines the prerequisite.
e03Approval denied for the requested changeThe prerequisite is not satisfied.
e04supplier_change committed; record version 18A side effect occurred despite the denial.
e05Read confirms the supplier at version 18Confirms the recorded business state.
e06Response says the update was approvedThe response conflicts with e03.

Move through an evidence pipeline with review gates

A useful analysis pipeline separates capture, checking, interpretation and release authority. The diagram follows the synthetic purchase-order example through a proposed sequence: select the review window, check completeness, prepare permitted evidence, apply defined checks, adjudicate uncertain findings, and evaluate a proposed change. Each stage produces an inspectable output rather than an implied guarantee.

At the completeness gate, PO-DEMO-042 keeps its approval, mutation and state-check references. At the checking gate, the reviewer receives approval_bypass as a proposed label with e03 and e04 attached. At adjudication, the reviewer confirms the policy applies to this action. A release gate is reached only after an owned change has been tested on separate tasks under agreed acceptance criteria.

Keep the evidence-missing path visible. A run without the decisive approval record can enter a repair or human-review queue instead of being forced into success or failure. A run with a known transport error may pass a deterministic check but still need outcome verification. The visual branches represent these decisions; their size and position do not represent dataset proportions.

Log the analysis version and each gate decision. If a later policy or taxonomy revision changes an interpretation, retain the previous review record and explicitly create a new one. The pipeline then provides a way to reproduce decisions across a collection. It is a recommended operating pattern, with no claim that a diagram alone creates an automated or complete control system.

From recorded events to an evaluated change

Follow the evidence through six review stages. A proposed finding is checked before it becomes a change decision.

  1. Stage 1

    Review window

    Input
    Recorded events and workflow versions

    Select the task family, time window and review question.

    Output
    A declared set of runs
  2. Stage 2

    Completeness

    Input
    Runs and required event families

    Check event identity, causal links and decisive records.

    Output
    Evidence status and gap references

    Missing decisive record? Repair capture or send for human review.

  3. Stage 3

    Permitted evidence

    Input
    Authorised records and access rules

    Keep decisive references; mark redactions and omissions.

    Output
    A compact, linked evidence view
  4. Stage 4

    Defined checks

    Input
    Evidence and versioned label definitions

    Apply a suitable rule or classifier; cite supporting events.

    Output
    Proposed labels with references

    Unsupported interpretation? Abstain and request review.

  5. Stage 5

    Adjudication

    Input
    Proposed labels, policy and event references

    Resolve exceptions and disagreements with an accountable reviewer.

    Output
    A reviewed finding or an unresolved case
  6. Stage 6

    Change evaluation

    Input
    An owned candidate change and separate test tasks

    Compare acceptance results with the baseline and record limits.

    Output
    Evidence for a release decision

    Acceptance unmet? Revise the change before seeking release approval.

An original proposed review method. Stage numbers show order, not volumes or performance. Missing evidence and unresolved interpretations remain visible; a label does not authorise release.

Reduce input without removing the decisive event

Preprocessing should make relevant evidence easier to inspect while preserving the basis of the finding. Build a compact operational view from stable event references, permitted excerpts, action identity, prerequisite results and side-effect outcomes. Keep the source record within its authorised retention boundary. Mark every redaction, omission and truncation so the reader can judge the view's limits.

For PO-DEMO-042, the compact view must retain e03's denied approval and e04's committed change. Removing a repeated catalogue response may be harmless to this question. Removing the approval because it looks like routine metadata changes the question itself. A summary saying that the order was updated would erase the evidence needed to distinguish an approved update from a bypass.

Lost in the Middle found that information position affected the tested retrieval and question-answering tasks. LongLLMLingua studies prompt compression in specified long-context tasks. These papers motivate testing how evidence selection affects your classifier; they do not establish that a particular current model will miss an event or that compression preserves every relevant detail.

Create a review set with long traces, late prerequisites, repeated tool outputs and contradictory messages. Compare labels and supporting references before and after reduction. Inspect disagreements and record whether the missing evidence was relevant. Version the reduction procedure with the rubric. A lower token bill is useful only when the resulting view still supports the decisions the team intends to make.

Sources: Lost in the Middle: How Language Models Use Long Contexts (external site)LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression (external site)

Create failure labels with boundaries and owners

A failure taxonomy is a versioned vocabulary for observed patterns. Start by reviewing varied runs with the people who own the workflow, then draft a small set of labels that answer operational questions. Give each label inclusion evidence, exclusions, severity and an owner. Keep observations, proposed causes and business consequences in separate fields.

The interactive taxonomy uses an original six-label procurement example. Selecting a label should reveal its meaning and the evidence needed to apply it. The table provides the same definitions in readable text. These categories are teaching examples, not a standard, a customer result or a released Bibha classifier. A run may carry several labels, while unknown means the current definitions cannot settle its interpretation.

For approval_bypass, require evidence that the applicable approval was denied or otherwise unsatisfied when the specific change committed. A missing approval event alone is excluded because the collection may be incomplete. evidence_missing records that limitation. This boundary stops a data-capture defect from being automatically reported as a policy violation and gives the instrumentation owner a separate issue to resolve.

Use discovery reviews to improve definitions before measuring a release. Record the policy and taxonomy version, keep difficult examples, and freeze the vocabulary for each comparison window. If a category later splits, document the mapping and re-score a common comparison set where needed. New labels can be useful, but their arrival should not be mistaken for a deterioration in the agent.

Original, illustrative multi-label taxonomy for procurement review. Severity descriptions concern the example and require workflow-specific review. No frequencies or benchmark scores are implied.
LabelInclude whenExclude whenIllustrative severityReview owner
approval_bypassA change commits despite an applicable unsatisfied approval.Approval evidence is absent or belongs to a different action.High for this denied supplier changeWorkflow and approval owner
unverified_stateA completion claim lacks the state evidence required by the acceptance contract.The required independent state check is present and agrees.Depends on the claimed outcomeIntegration owner
duplicate_actionTwo committed side effects exceed the permitted single action.A retry is rejected or the operation is safely idempotent.High if the extra side effect is consequentialIntegration owner
evidence_missingRequired records are missing, truncated or unavailable.All evidence needed for this finding is available.Evidence limitation; operational risk unresolvedInstrumentation owner
tool_errorA recorded tool operation returns an error or uncertain transport outcome.Only the business outcome is wrong with no recorded tool error.Depends on recovery and final stateTool owner
unknownAvailable evidence cannot be assigned reliably under this taxonomy.A defined label is supported or required evidence is missing.Review required before assigning riskDomain reviewer

Which finding does the evidence support?

Expand an example to compare its definition, boundary and observed events. A run can support several labels; these cases isolate different decisions.

Synthetic teaching examples, not customer data or a benchmark

approval_bypassA denied change still commits
Definition
A change commits despite an applicable approval requirement being unsatisfied.
Include when
The denial and committed change refer to the same action under the applicable policy, with no valid exception.
Exclude when
Approval evidence is missing, applies to another action, or a valid policy exception permits the change.

Synthetic trace: PO-DEMO-042

  1. e03P-7 approval denied for the requested supplier change.
  2. e04That supplier change commits at order version 18.
  3. e05Independent read-back verifies the changed supplier at version 18.
Supported finding
approval_bypass only. The read-back verifies state, so this example does not support unverified_state.
What to review
Confirm policy P-7, action identity and causal order; inspect where the workflow should enforce the approval requirement.
unverified_stateCompletion is claimed before state is checked
Definition
A completion claim lacks the state evidence required by the acceptance contract.
Include when
A fully captured workflow claims completion but does not perform the required independent state check.
Exclude when
The required state check is present and agrees with the completion claim. An accepted request alone is not a state check.

Synthetic trace: UPDATE-DEMO-017

  1. e01Write request accepted for asynchronous processing.
  2. e02Workflow ends without invoking the required read-back; capture is complete.
  3. e03Final response says the record was updated.
Supported finding
unverified_state. The final business state remains unresolved. The verification step was omitted, rather than lost by the collector.
What to review
Confirm the acceptance contract and run completeness; obtain a permitted state check before treating the update as complete.
duplicate_actionA retry creates a second side effect
Definition
Two committed side effects exceed the permitted single action.
Include when
Distinct operation results confirm two committed effects for one action that permits only one.
Exclude when
The retry is rejected, reuses an idempotent result, or the collector records the same event twice.

Synthetic trace: RESERVE-DEMO-016

  1. e01Workflow authorises one reservation.
  2. e02First operation commits reservation R-DEMO-201.
  3. e03Retry commits a distinct reservation R-DEMO-202.
Supported finding
duplicate_action. Two distinct committed results support the finding; the number of attempts alone would not.
What to review
Confirm the single-action constraint, operation identities and resulting records; inspect idempotency and retry handling.
evidence_missingThe decisive approval record is unavailable
Definition
Required evidence is missing, truncated or unavailable for the review question.
Include when
The collector or permitted evidence view cannot supply the record needed to establish the approval decision.
Exclude when
All decisive records are available. An adverse outcome does not by itself establish a capture gap.

Synthetic trace: CAPTURE-DEMO-023

  1. e01Collector marks the required approval record as lost.
  2. e02Supplier change is recorded as committed.
  3. e03Read-back confirms the changed supplier.
Supported finding
evidence_missing. The mutation is verified, but the trace cannot establish whether approval was granted or denied.
What to review
Recover authorised approval evidence and inspect collector loss; keep approval_bypass undecided until its prerequisite evidence exists.
tool_errorA timeout leaves the write outcome uncertain
Definition
A tool operation records an error or an uncertain transport outcome.
Include when
The tool records a definite failure, or a transport timeout whose business outcome has not yet been resolved.
Exclude when
Only the business outcome is wrong and no tool error is recorded. A verified later success does not erase a recorded transport error.

Synthetic trace: WRITE-DEMO-031

  1. e01Write request is sent.
  2. e02Transport times out without a commit result.
  3. e03Run closes with write outcome explicitly unresolved; capture is complete.
Supported finding
tool_error with an uncertain outcome. This timeout does not prove rejection or commit; a definite validation rejection would prove that attempt failed.
What to review
Inspect operation identity and retry safety, then perform a permitted state check. Do not retry blindly when the first write may have committed.
unknownComplete evidence reaches an undefined exception
Definition
Available evidence cannot be assigned reliably under the current taxonomy.
Include when
The decisive records are available, but an unfamiliar policy interpretation is outside the defined label boundaries.
Exclude when
A defined label is supported, or missing decisive records prevent interpretation. Missing records belong under evidence_missing.

Synthetic trace: EXCEPTION-DEMO-008

  1. e01Policy returns delegated_exception with its complete decision record.
  2. e02The scoped change commits and read-back confirms the resulting state.
  3. e03Collector confirms that decision, write and state-check records are complete.
Supported finding
unknown pending domain review: the rubric does not define delegated_exception. The gap is in interpretation, not in the recorded evidence; bypass is not established.
What to review
Ask the policy owner to resolve the exception, document the decision and version the taxonomy before comparing later results.
An original illustrative taxonomy. Event IDs identify fictional records, not observed frequencies. A trace records observable execution; it does not reveal private internal reasoning.

Match the check to the available evidence

Use deterministic checks when the relevant condition is explicit in structured data. Use semantic classification when interpretation depends on policy language or several related events. Use human adjudication for disputed, unfamiliar or consequential cases. These methods can work together; none should inherit execution authority simply because it produces a label.

A deterministic rule can compare e03's decision with e04's commit state if their action IDs and policy scope are reliable. An LLM-based semantic classifier can propose that e06 contradicts the recorded approval, returning the supporting event IDs and a short explanation. A domain reviewer can resolve whether an exceptional approval path applies. Require the same evidence discipline from all three.

Model-based judges need their own checks. The MT-Bench judge study identifies position, verbosity and self-enhancement biases in its evaluated settings. For this workflow, test whether reordering equivalent evidence, changing response length or naming the model alters labels. Those experiments establish limits for the chosen judge and rubric instead of borrowing another benchmark's reported agreement.

Keep classifier input and output constrained. Supply the permitted evidence, applicable policy and current definitions; validate returned label IDs and event references. An explanation without a valid reference is a review suggestion, not an established finding. When evidence is unavailable or interpretation conflicts, let the checker abstain and route the case to the appropriate queue.

Choose the least complex check that can answer the review question.
MethodUseful forLimitRequired output
DeterministicExplicit denial plus matching committed action; duplicate IDs; missing fieldsOnly as sound as the recorded fields and rule scopeRule version, result and event references
SemanticPolicy interpretation, contradictions and cross-event meaningCan misread evidence or produce unsupported explanationsRubric version, proposed labels, references and abstention
HumanExceptions, disagreement, new patterns and consequential findingsReview capacity and consistent guidance are requiredDecision, rationale, evidence and accountable reviewer

Sources: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (external site)

Build reviewed examples before measuring a classifier

Create a reference set that represents the decisions you want the classifier to make. Include accepted runs, relevant failures, incomplete evidence and legitimate exceptions. Ask domain reviewers to label the evidence independently where practical, then adjudicate disagreements. Their resolved decisions become the reference for that dataset, with limits recorded rather than described as infallible truth.

Separate label discovery, development, threshold tuning and the final test. Keep related attempts, duplicate tasks and events from one run in the same split. Where release generalisation matters, reserve a later time window or a task family that development has not touched. The split should reflect the future decision, not merely make the spreadsheet convenient.

For procurement, include approval granted, approval denied, approval expired, a permitted exception, a timed-out write with confirmed state and a timed-out write with unresolved state. Add a duplicate request whose idempotency key prevents a second mutation. These cases check whether the definitions distinguish superficially similar events with different operational meanings.

Hide reference diagnoses and expected labels from the classifier's inputs. Keep a record of how examples were selected and which cases lack a settled reference. If the team changes a label definition after reading final-test mistakes, that set has become development material. Reserve fresh evaluation evidence before claiming the revised system generalises.

Report precision, recall and calibrated decisions separately

Measure each important label, including its support and review burden. Precision asks how many proposed positives are correct; recall asks how many reference positives were found. For an explicitly synthetic result with 80 true positives, 20 false positives and 40 false negatives, precision is 80% and recall is about 66.7%. These are arithmetic examples, not measured classifier results. Apply these checks to both Jev decisions and LLM-generated labels.

Report the confusion counts, dataset window and per-label sample size beside those ratios. A rare approval bypass can disappear inside overall accuracy when most runs are acceptable. Multi-label evaluation also needs a stated counting unit. If no relevant positives or predictions exist, report the score as undefined where appropriate and explain the available evidence instead of inventing a perfect value.

Calibration is a separate property. It compares predicted probabilities with observed label frequencies; a model's written confidence score is not automatically a probability. Fit or assess calibration using appropriate separate data, and tune action thresholds away from the untouched final test. The scikit-learn documentation distinguishes probability estimation, calibration and threshold choice.

Choose thresholds according to the action. An investigation queue may accept more false alarms to find additional cases, while a consequential operational decision requires stricter evidence and review. Show what each threshold does to missed cases and reviewer capacity. Recheck performance when the model, rubric, policy, preprocessing or workload changes; the previous measurement describes its recorded configuration.

Sources: Precision, recall, F-score and support | scikit-learn (external site)Probability calibration | scikit-learn (external site)Tuning the decision threshold for class prediction | scikit-learn (external site)

Sample for coverage and keep uncertainty visible

Separate retention sampling from investigation sampling and population measurement. OpenTelemetry describes head sampling as an early decision and tail sampling as a decision using later trace evidence. Its tail-sampler implementation has routing, memory and late-span constraints. A configured error policy cannot establish that every failing run was retained, especially if evidence was dropped upstream.

Maintain a probability-sampled baseline where feasible, alongside targeted cohorts for rare events, new versions, slow runs and incomplete capture. Record why each run was selected and its effective inclusion probability when that probability is valid. OpenTelemetry's probability-sampling guidance defines adjusted counts under stated assumptions; targeted selections without known probabilities do not acquire valid weights merely because they are useful for debugging.

A collection deliberately enriched for approval errors is suitable for investigating those errors. Its raw label percentage is not the production failure rate. Population estimates need the correct denominator, valid selection weights, duplicate handling and consideration of classification errors. Report the target population and the uncertainty around the estimate rather than presenting a precise number without a defensible design.

Keep unknown and evidence_missing visible in reporting. Count their review queues, ageing and resolution separately from confirmed failures and accepted outcomes. Do not silently call abstention success. In multi-label reporting, label counts can exceed run counts, so show both. Rising unknowns may indicate a new workflow, a weak definition or changed evidence capture; review examples before selecting the remedy.

A trace record branches into three failure-category boxes, with a separate dashed path to an amber uncertain-review box.
Original AI-generated teaching diagram of failure classification and a separate route for cases that need human review.

Sources: Sampling | OpenTelemetry (external site)Tail Sampling Processor | OpenTelemetry Collector Contrib (external site)TraceState: Probability Sampling | OpenTelemetry (external site)

Estimate token cost and the work around it

Estimate cost from the actual analysis plan, including selection and repeated passes. Let N be eligible retained traces, s the selection fraction and k the passes per selected trace. Multiply N × s × k by average input and output tokens per pass. Price non-overlapping billed categories using the same currency and analysis window.

Here is a wholly illustrative example: 100,000 traces × 0.25 selected × 2 passes equals 50,000 passes. At 1,500 input and 100 output tokens per pass, that is 75 million input tokens and 5 million output tokens. Assume 1 cost unit per million input tokens and 4 per million output tokens. Classification then costs 75 × 1 + 5 × 4 = 95 units.

These volumes and prices are assumptions, not Jev pricing, provider tariffs, Bibha results or a quote. The example excludes retries, caching, request rounding and differing trace lengths until those terms are added. Google publishes distinct billing categories for its models and modes; consult the actual provider contract when building an estimate. Avoid adding cached-input or reasoning-detail counters again when they are already included in billed totals.

Add discovery work, data preparation, storage, scanning, deterministic compute, orchestration and human review separately. Test sensitivity to longer evidence, more passes and a larger selected fraction. Then ask whether the reduced budget still covers important failures and uncertain cases. A cheap classifier that overwhelms reviewers or omits the decisive event can make the complete operating process more expensive.

Assumed classification cost only. All prices are illustrative cost units.
TermCalculationResult
Selected traces100,000 × 0.2525,000
Classifier passes25,000 × 250,000
Input tokens50,000 × 1,50075,000,000
Output tokens50,000 × 1005,000,000
Input cost75 × 175 units
Output cost5 × 420 units
Classification total75 + 2095 units before other costs
Inspectable cost equations. Use averages and prices from the same analysis window.
C_classification = N * s * k * ((t_in / 1000000) * p_in + (t_out / 1000000) * p_out)
C_total = C_classification + C_discovery + C_data_and_storage + C_deterministic_compute + C_orchestration + C_human_review

Sources: Gemini Developer API pricing | Google AI for Developers (external site)

Treat trace content as untrusted evaluator input

Protect the analysis system as carefully as the agent being reviewed. Trace bodies can contain customer text, retrieved documents and tool responses that an attacker influences. AgentDojo evaluates prompt-injection attempts delivered through untrusted tool content. This supports testing the evaluator's trust boundary separately from ordinary task-quality checks; it does not establish an immune defence.

For the proposed procurement review, give the evaluator read-only access to the permitted evidence view and no authority to change orders. Treat embedded requests to ignore policy or reveal data as content to analyse. Keep evaluator instructions, policy and source evidence clearly separated. Validate the output against allowed labels and verify that cited events belong to the authorised run.

Minimise content capture and agree access, processing location, retention and deletion with the system owner. Redact secrets and personal details where required before export, with redaction marked in the evidence view. W3C Trace Context forbids personal or sensitive information in correlation headers. Trace IDs should connect records; they should not serve as permission to retrieve them.

Test fabricated event references, malicious retrieved text, cross-organisation IDs and attempts to trigger evaluator tools. Include incomplete and redacted traces in those tests. Record utility and boundary failures independently so an apparently helpful diagnosis cannot conceal an access violation. Keep logs free of unnecessary source bodies and provide a controlled route for authorised reviewers to inspect the underlying evidence.

Sources: AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents (external site)Trace Context | W3C Recommendation (external site)

Turn a reviewed finding into a controlled release

Use trace findings to justify an owned change and a repeatable acceptance test. For PO-DEMO-042, the first task is to confirm where the approval requirement should be enforced, then test the smallest relevant change. A proposed label does not authorise a production mutation. The workflow owner and release authority remain responsible for the decision.

Evaluate the complete task, including final state. tau-bench demonstrates checking a post-interaction database state against an annotated goal and assessing repeated success. Apply that idea carefully to your own acceptance contract: a denied approval should leave the supplier unchanged, an approved change should reach the intended state, and a retry should not create a duplicate side effect.

Compare the baseline and candidate on a fixed set, a held-out set and the required adversarial cases. Keep task mix, policy, taxonomy and evaluator versions constant where the comparison needs them. Report quality, latency, resource use and unresolved cases together. If the measurement procedure changes, separate that change from the system improvement you are trying to establish.

Before release, name the approver, rollout scope, alert owner and rollback trigger. Check recovery from a partial rollout or failed dependency. After release, inspect fresh runs, the sampled baseline and unresolved queues under the same declared contract. Record what changed, the evidence supporting it and what remains uncertain. A controlled release closes one review cycle while preserving the means to discover the next issue.

  • A practical release record. Record the review window and versions, approved evidence access, adjudicated findings, change owner, acceptance results, remaining limitations, release authority and rollback plan. Attach the permitted evidence references so the next reviewer can assess the decision.

Sources: tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (external site)

Connect the evidence to the workflow you operate

Bibha helps teams connect task traces and operating measures with evaluation, reviewed feedback and controlled improvement releases. Its monitoring, evaluation and governance offering covers agreed business outcomes, task-specific tests, human review, versioned results and defined access and release responsibilities. The public platform pages explain that scope; implementation details depend on the system and engagement.

The event contract, storage choices, taxonomy and classifier pipeline in this guide are educational recommendations. They should be assessed against your workflow, evidence permissions and operating constraints before becoming a design. No Bibha throughput, classifier score, saving or automatic production-learning result is claimed here. Monitoring evidence helps people make a reviewed change; it does not mean every interaction changes the live model.

Prepare a workflow brief with the decision you want to improve, the systems involved, the authoritative outcome and the approval owner. Add the evidence you are permitted to retain, the review volume and the operating budget. Those inputs make it possible to discuss the actual boundary between platform configuration, integration work and ongoing review.

If you are planning an agent workflow, bring that brief to a conversation with Bibha. Start with one consequential question, such as whether a change can occur without current approval, and agree how its evidence will be checked. A useful first outcome is a scoped review and release plan with named responsibilities and acceptance criteria.

Questions and answers

Is Jev a Bibha or Applied Compute model?

Jev is a model from TypeSafe AI. This guide assesses its possible role in trace review without claiming Bibha owns Jev or provides an approved Jev integration.

Does Jev confidence prove that a trace label is correct?

No. For Choice and Score, confidence is derived from the returned probability distribution. Noul has no separate confidence field. Test the questions, evidence and version on reviewed examples and retain appropriate human adjudication.

How do traces differ from ordinary logs?

Logs record individual messages or events. A trace links recorded operations within an execution path. For business review, connect those technical records to the task and its authoritative outcome. Both can contribute evidence; neither is complete simply because it exists.

Can an LLM classify every failure reliably?

No reliability claim follows from using an LLM. Evaluate the chosen model, evidence view and rubric on reviewed, held-out examples. Report per-label errors and abstentions, and retain human review for uncertain or consequential findings.

What should happen when an approval event is missing?

Mark the evidence gap and check permitted source records or capture health. A missing approval event does not prove denial or approval. In the procurement example, reserve approval_bypass for a supported unsatisfied prerequisite and committed matching action.

Is a confidence score a calibrated probability?

Only if appropriate evaluation supports that interpretation. A generated confidence number alone is insufficient. Assess its relationship to observed label frequencies on separate data, and choose action thresholds according to the decision and its consequences.

Can a targeted sample show the production failure rate?

Its raw percentage cannot establish that rate. Targeted incident samples deliberately change the mix. Population estimates require a defined population, suitable inclusion probabilities or sampling design, defensible denominators and consideration of classifier error.

How can we lower analysis cost without losing evidence?

Query metadata first, select relevant permitted events, remove repetition carefully and measure the effect on findings. Count passes and retries, then add storage, compute and review costs. Preserve decisive evidence and mark every omission.

Does analysing traces reveal hidden model reasoning?

No. This guide uses instrumented operations, recorded content and observable state. Explanations about causes remain hypotheses until tested; a trace does not provide access to private internal reasoning.

Should labels automatically change a production agent?

A label is analysis output. Turn it into an owned issue, verify the cause, test a proposed change and obtain the accountable release decision. Keep release authority separate from the diagnostic evaluator.

References

  1. Billion-Token Scale Trace Analysis: Jev vs LLMs | Applied Compute https://www.appliedcompute.com/platform/billion-token-scale-trace-analysis (external site)
  2. Traces | OpenTelemetry https://opentelemetry.io/docs/concepts/signals/traces/ (external site)
  3. Trace Context | W3C Recommendation https://www.w3.org/TR/trace-context/ (external site)
  4. Semantic conventions for generative AI systems | OpenTelemetry https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/README.md (external site)
  5. Semantic conventions for generative client AI spans | OpenTelemetry https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md (external site)
  6. Sampling | OpenTelemetry https://opentelemetry.io/docs/concepts/sampling/ (external site)
  7. Tail Sampling Processor | OpenTelemetry Collector Contrib https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/processor/tailsamplingprocessor/README.md (external site)
  8. TraceState: Probability Sampling | OpenTelemetry https://opentelemetry.io/docs/specs/otel/trace/tracestate-probability-sampling/ (external site)
  9. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena https://arxiv.org/abs/2306.05685 (external site)
  10. Precision, recall, F-score and support | scikit-learn https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_recall_fscore_support.html (external site)
  11. Probability calibration | scikit-learn https://scikit-learn.org/stable/modules/calibration.html (external site)
  12. Tuning the decision threshold for class prediction | scikit-learn https://scikit-learn.org/stable/modules/classification_threshold.html (external site)
  13. Lost in the Middle: How Language Models Use Long Contexts https://arxiv.org/abs/2307.03172 (external site)
  14. LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression https://arxiv.org/abs/2310.06839 (external site)
  15. tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains https://arxiv.org/abs/2406.12045 (external site)
  16. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents https://arxiv.org/abs/2406.13352 (external site)
  17. Reading and Writing Parquet Files | DuckDB https://duckdb.org/docs/current/data/parquet/overview (external site)
  18. Hive Partitioning | DuckDB https://duckdb.org/docs/current/data/partitioning/hive_partitioning (external site)
  19. Gemini Developer API pricing | Google AI for Developers https://ai.google.dev/gemini-api/docs/pricing (external site)
  20. TypeSafe AI: introducing System One models and Jev https://typesafe.ai/blog/introducing-system-one-models-and-jev (external site)
  21. TypeSafe AI: Jev models, versions and limits https://docs.typesafe.ai/models (external site)
  22. TypeSafe AI: Choice, Score and Noul questions https://docs.typesafe.ai/primitives (external site)
  23. TypeSafe AI: confidence and probability distributions https://docs.typesafe.ai/confidence (external site)
Explore AI agent configurationExplore monitoring and improvementExplore evaluation and release qualityReview governance and data boundariesPrepare an AI system for operationDiscuss your agent workflowAll news and research