aeat.agent.eval package¶
Operator golden-task eval: trajectory + provenance assertions over the harness.
The operator-side counterpart of the persona testimonial gate. A golden scenario declares the expected tool trajectory for a workflow and the provenance the result must carry; the runner asserts the trajectory resolves against the live CLI surface, follows the modelo lifecycle order, is consistent with the shipped skill playbook, and that the modelo’s casillas carry their legal grounding in the registry. It reuses the scenario methodology of the persona testimonials without inheriting their knowledge-withholding brief, and it never hand-computes a tax value (that would be a tautological calculation test).
Submodules¶
- aeat.agent.eval._flywheel module
- aeat.agent.eval._live_harness module
- aeat.agent.eval._live_scoring module
FaithfulnessCheckFnLiveScenarioScoreLiveScenarioScore.scenarioLiveScenarioScore.personaLiveScenarioScore.session_idLiveScenarioScore.keys_resolveLiveScenarioScore.lifecycle_orderedLiveScenarioScore.expected_coveredLiveScenarioScore.tool_errorsLiveScenarioScore.invariantsLiveScenarioScore.narration_checksLiveScenarioScore.failuresLiveScenarioScore.passed
score_live_trajectory()DiscoveryScoreSurfaceDiscoveryComparisonSurfaceDiscoveryComparison.target_command_keySurfaceDiscoveryComparison.coreSurfaceDiscoveryComparison.fullSurfaceDiscoveryComparison.core_advertised_tool_countSurfaceDiscoveryComparison.full_advertised_tool_countSurfaceDiscoveryComparison.failuresSurfaceDiscoveryComparison.both_reached_same_targetSurfaceDiscoveryComparison.rounds_deltaSurfaceDiscoveryComparison.full_surface_advertises_moreSurfaceDiscoveryComparison.passed
score_discovery_trajectory()compare_surface_discovery()IdentityStateProtocolIdentityGateRefusalFnIdentityConfirmationScoreIdentityConfirmationScore.scenarioIdentityConfirmationScore.personaIdentityConfirmationScore.session_idIdentityConfirmationScore.mutating_step_presentIdentityConfirmationScore.identity_confirmedIdentityConfirmationScore.gate_refused_mutationsIdentityConfirmationScore.failuresIdentityConfirmationScore.passed
score_identity_trajectory()
- aeat.agent.eval._models module
GoldenScenarioNarrationFaithfulnessConfirmationTierConfirmationGateCheckGoldenResultGoldenResult.scenarioGoldenResult.trajectory_resolvesGoldenResult.lifecycle_orderedGoldenResult.skill_consistentGoldenResult.provenance_presentGoldenResult.response_provenance_presentGoldenResult.verification_groundedGoldenResult.narration_faithfulness_checksGoldenResult.expected_confirmation_tiersGoldenResult.failuresGoldenResult.passed
ExitCodeScenarioExitCodeVerdictUnderDeclarationScenarioUnderDeclarationVerdictContradictionScenarioContradictionVerdictProfileConfirmationScenarioProfileConfirmationVerdictElicitationActionLiveToolCallRecordLiveNarrationRecordLiveElicitationRecordLiveTrajectoryLiveInvariantVerdict
- aeat.agent.eval._replay module
- aeat.agent.eval._report module
ScenarioOutcomeRowMeasurementReportMeasurementReport.scenarios_runMeasurementReport.scenarios_passedMeasurementReport.live_submit_attempts_totalMeasurementReport.handoff_faithfulness_blocks_totalMeasurementReport.tool_errors_totalMeasurementReport.unfaithful_narrations_totalMeasurementReport.rowsMeasurementReport.invariants_holdMeasurementReport.all_passed
build_measurement_report()render_measurement_report_markdown()
- aeat.agent.eval._runner module