MYTHOS-AI-0003
Applied Natural Language and LLM Evaluation Standard
Evaluation standard for applied natural-language tasks grounded in retrieval, citation, and operational correctness.
Purpose
Provide a method for evaluating applied LLM behavior on Mythos public and operational corpora, including retrieval grounding and refusal quality.
Scope
Applies to Ask Mythos, guide surfaces, and other NL interfaces that answer from approved content or structured tools.
Does not invent accuracy percentages without a named suite.
Evaluation dimensions
Grounding: answers cite approved sources when required.
Hallucination and unsupported claim rate on held-out prompts.
Refusal quality for out-of-scope or prohibited requests.
Latency class reported only with hardware and model identifiers when published.
Publication rules
Published LLM evaluations must include prompt suite identity, corpus revision, model identity, date, and scoring rubric.
Comparative claims against other vendors require the same suite and disclosure rules.
Document status
MYTHOS-AI-0003 is published here as a draft standard definition. It does not constitute a completed product certification or a published scoreboard of model results.