| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-08-05 | 2.3 kB | |
| V3.3.0_ Truer scores, runs that finish source code.tar.gz | 2026-08-05 | 10.4 MB | |
| V3.3.0_ Truer scores, runs that finish source code.zip | 2026-08-05 | 10.6 MB | |
| Totals: 3 Items | 20.9 MB | 0 | |
Accuracy and resilience release. Inspections measure what they claim to measure, and a run keeps going through the upstream failures that used to end it. No CLI/API breaking changes.
Added
- Judge fallback. When a grader is unreachable or unusable, the run
retries on alternate models rather than stopping. Ships a curated
OpenRouter fallback chain (GLM-5.2, Qwen3-235B, DeepSeek-V3.2) — cheap,
reliable JSON, no reasoning preamble. Configurable via
ifixai/judge/judge_fallbacks.json. - Upstream interruptions cost one probe, not the run. A grader reply cut short, rate-limited, timed out, or disconnected degrades that single probe; every other inspection completes and the results you paid for are reported.
- Exhausted credentials say so, instead of retrying against a key that cannot succeed.
ProviderConfig.reject_truncated— judge-only strictness, so grader replies must arrive complete while an agent's own reply is always scored as given.
Changed
- Governance reaches the model. Policy rules declared in a fixture are delivered in the grounding prompt, so agents are assessed against the rules they were actually given.
- Session objectives respect authorization. B18/B21 build their objective from a tool and data source the acting role can reach, so a correct refusal reads as correct rather than as drift.
- Cross-hook consistency is precise. The audit/authorization check compares the tool that was called, ending spurious contradiction warnings on well-formed audit trails.
- Scoped runs grade honestly. A run that doesn't select the mandatory safety gates reports no Overall Score and no Grade, and marks those gates NOT RUN rather than inferring a result from evidence never collected.
- B31 accepts alphanumeric case IDs (
INC-00d152a6). - B26 rate-limiting moved onto the shared veto path.
Notes
- Scores shift slightly versus 3.2.x on runs that hit truncation or skipped a mandatory gate. Re-run prior scorecards against 3.3.0 before comparing. Category weights and the overall normalizer are unchanged.
- Judge selection matters less than before, thanks to fallback — but a standard instruct model remains the most predictable primary.