Guides and Field Guides / AI and Agentic Systems Field Guides
AI Safety Alignment Red Teaming Sycophancy and Control
A practical-to-frontier safety guide covering risk governance, objective and instruction alignment, sycophancy, reward hacking, deceptive behavior, prompt injection, tool misuse, memory poisoning, red teaming, control architectures, human oversight, monitoring, and incident response.
Publication record
MYTHOS-FG-AI-0012
- Publication class
- Focused Field Guide
- Status
- Candidate
- Pages
- 55
- File size
- 236 KB
Reader
Read the publication
Searchable, bookmarked PDF, 55 pages. Loads on request (236 KB).
Candidate status describes publication review state. It does not establish certification, legal compliance, implementation conformance, benchmark reproduction, or product readiness.
Related
More in AI and Agentic Systems Field Guides
MYTHOS-FG-AI-0001 AI Model Architecture and Capability Foundations
A foundation-to-frontier guide covering tensors, representation learning, transformers, attention, encoders, decoders, mixture-of-experts, diffusion, state-space and recurrent alternatives, multimodality, scaling, capability boundaries, and architecture evaluation.
MYTHOS-FG-AI-0002 AI Model and Agent Evaluation Engineering
A rigorous evaluation guide covering capability, quality, safety, robustness, calibration, agents, tools, long-horizon tasks, datasets, contamination, judges, human studies, performance, cost, regression, and decision-focused reporting.
MYTHOS-FG-AI-0003 Local Models On Device Inference and Sovereign AI
A local-first AI guide covering model acquisition, licensing, hardware sizing, GGUF, runtimes, accelerators, offline operation, privacy, sovereign controls, model lifecycle, evaluation, updates, and hybrid routing.
MYTHOS-FG-AI-0004 Inference Serving Quantization KV Cache and Throughput Engineering
A production inference guide covering serving architecture, tokenization, prefill, decoding, batching, scheduling, KV caches, quantization, parallelism, speculative decoding, autoscaling, SLOs, observability, cost, and failure recovery.
MYTHOS-FG-AI-0005 Fine Tuning Alignment Synthetic Data and Training Governance
A training and post-training guide covering objectives, datasets, SFT, PEFT, LoRA, QLoRA, preferences, DPO, reward models, synthetic data, curriculum, distributed training, evaluation, safety, provenance, release, and rollback.
MYTHOS-FG-AI-0006 Multimodal AI Vision Voice and Realtime Interaction
A multimodal systems guide covering image, document, audio, speech, video, fusion, grounding, realtime streaming, voice activity, turn-taking, latency, accessibility, privacy, safety, and multimodal evaluation.