MYTHOS-AI-0004
AI Software Engineer and Code Generation Benchmark Standard
Benchmark standard for code generation, repair, and review assistants under typed constraints and repository context.
Purpose
Define reproducible evaluation for AI-assisted software engineering: generation, repair, review, and test synthesis under repository constraints.
Scope
Applies to coding agents and assistants that edit or propose code inside Mythos engineering workflows.
Language coverage and repository fixtures must be declared per published run.
Evaluation dimensions
Functional correctness against fixture tests.
Constraint adherence: typing, style, and security lint gates.
Context use: relevant file selection versus noise.
Human review burden: clarity of diffs and rationale.
Publication rules
Published code benchmarks must name languages, fixture repos, model or agent version, date, and pass criteria.
No pass-rate claims without the suite identity.
Document status
MYTHOS-AI-0004 is published here as a draft standard definition. It does not constitute a completed product certification or a published scoreboard of model results.