MYTHOS-AI-0004

AI Software Engineer and Code Generation Benchmark Standard

Benchmark standard for code generation, repair, and review assistants under typed constraints and repository context.

Draft document Artificial Intelligence & Agentic Systems

Purpose

Define reproducible evaluation for AI-assisted software engineering: generation, repair, review, and test synthesis under repository constraints.

Scope

Applies to coding agents and assistants that edit or propose code inside Mythos engineering workflows.

Language coverage and repository fixtures must be declared per published run.

Evaluation dimensions

Functional correctness against fixture tests.

Constraint adherence: typing, style, and security lint gates.

Context use: relevant file selection versus noise.

Human review burden: clarity of diffs and rationale.

Publication rules

Published code benchmarks must name languages, fixture repos, model or agent version, date, and pass criteria.

No pass-rate claims without the suite identity.

Document status

MYTHOS-AI-0004 is published here as a draft standard definition. It does not constitute a completed product certification or a published scoreboard of model results.