MYTHOS-AI-0003

Applied Natural Language and LLM Evaluation Standard

Evaluation standard for applied natural-language tasks grounded in retrieval, citation, and operational correctness.

Draft document Artificial Intelligence & Agentic Systems

Purpose

Provide a method for evaluating applied LLM behavior on Mythos public and operational corpora, including retrieval grounding and refusal quality.

Scope

Applies to Ask Mythos, guide surfaces, and other NL interfaces that answer from approved content or structured tools.

Does not invent accuracy percentages without a named suite.

Evaluation dimensions

Grounding: answers cite approved sources when required.

Hallucination and unsupported claim rate on held-out prompts.

Refusal quality for out-of-scope or prohibited requests.

Latency class reported only with hardware and model identifiers when published.

Publication rules

Published LLM evaluations must include prompt suite identity, corpus revision, model identity, date, and scoring rubric.

Comparative claims against other vendors require the same suite and disclosure rules.

Document status

MYTHOS-AI-0003 is published here as a draft standard definition. It does not constitute a completed product certification or a published scoreboard of model results.