AI Safety, Governance & Alignment Practice (FIELD GUIDE (ALL LEVELS))
Start this planThe central quality signal the field guide holds teams accountable to — defining what 'good enough' output means for safety-critical use and orchestrating the levers (data, prompts, retrieval, adaptation) that move it across the portfolio. (Setting the org's safety standard and defending it as a field guide.)
started 0 · finished 0 (claimed) · not yet measured (verified) · no data (n<5)
Ordered tasks (43) — this is what a project auto-creates
- #1Model and Provider SelectionHighTo Do
Use the criteria and escalation frame for choosing models, providers, and architectures across a portfolio of safety-critical systems — not for a single build.
Assignee: Unassigned · due 0 days after project start · takes 14 days
- #2Training Data Quality, Coverage, and QuantityHighTo Do
Define the data-quality standard you hold across teams — treating provenance, coverage, and bias exposure as governance obligations rather than engineering nice-to-haves.
Assignee: Unassigned · due 14 days after project start · takes 14 days
- #3Hallucination RateHighTo Do
Turn 'the model makes things up' into a governed metric with risk-tiered thresholds and known reduction levers.
Assignee: Unassigned · due 28 days after project start · takes 14 days
- #4Prompt Engineering QualityHighTo Do
Use the practices for making prompt design a reviewed, versioned discipline rather than individual craft that varies by author.
Assignee: Unassigned · due 42 days after project start · takes 14 days
- #5Data Warehouse / System Architecture DesignHighTo Do
Set the structural foundation — storage, partitioning, replication, and access patterns — beneath safety-critical models. You get the case for making architecture an org-wide standard rather than a per-project choice.
Assignee: Unassigned · due 56 days after project start · takes 14 days
- #6Data Integration and ETL QualityHighTo Do
Address the rigor of the pipelines that combine heterogeneous sources into the data your models consume. You learn why pipeline integrity is a precondition for any credible safety claim.
Assignee: Unassigned · due 70 days after project start · takes 14 days
- #7Data Quality and Information TrustHighTo Do
Owning the accuracy, completeness, currency, and conformance of your delivered data — and the confidence stakeholders reasonably place in it. You get the case for treating trust as a governed, measured property.
Assignee: Unassigned · due 84 days after project start · takes 14 days
- #8Requirements and Business Objective AlignmentMediumTo Do
Anchor safety work to genuine problems and purposes rather than to whatever the available data and tooling make convenient. You get the practice of defining the problem before designing the solution.
Assignee: Unassigned · due 98 days after project start · takes 7 days
- #9Fine-Tuning and Model AdaptationHighTo Do
Judge when adapting a pretrained model genuinely improves safety and when it silently degrades alignment already baked into the base model.
Assignee: Unassigned · due 105 days after project start · takes 14 days
- #10Observability and Monitoring MaturityHighTo Do
Define what 'observable enough to be safe' means and how you ensure teams instrument systems to that bar before shipping, not after an incident.
Assignee: Unassigned · due 119 days after project start · takes 14 days
- #11Generalization / Model Variance ControlHighTo Do
Equips you to reason about and communicate out-of-distribution robustness — distinguishing models that genuinely generalize from those that only look strong on the test set.
Assignee: Unassigned · due 133 days after project start · takes 14 days
- #12Deployment and Serving ArchitectureMediumTo Do
Treats deployment and serving patterns as a safety boundary — architecting rollout gates, rollback paths, and serving robustness so unsafe behavior is contained before it reaches users.
Assignee: Unassigned · due 147 days after project start · takes 7 days
- #13Inference Efficiency and CostMediumTo Do
Hold the line where cost and latency pressure meets safety margin — deciding when efficiency gains must not be allowed to thin out monitoring, guardrails, or evaluation.
Assignee: Unassigned · due 154 days after project start · takes 7 days
- #14Guardrails and AI Security ControlsHighTo Do
Define the guardrail and security-control regime that constrains unsafe model behavior in production, and how you get it implemented consistently instead of team-by-team improvisation.
Assignee: Unassigned · due 161 days after project start · takes 14 days
- #15Data Distribution Shift and Model FreshnessHighTo Do
Drift and staleness as an ongoing threat you must detect and respond to on a horizon longer than any single release.
Assignee: Unassigned · due 175 days after project start · takes 14 days
- #16Context Grounding / Retrieval RelevanceHighTo Do
Sets the standard for how tightly outputs must be tied to retrieved evidence before they can be trusted in safety-critical use.
Assignee: Unassigned · due 189 days after project start · takes 14 days
- #17Metadata Management and Data GovernanceMediumTo Do
Establish lineage, metadata, and governance so data provenance is traceable and auditable across the org. You learn why this is a foundational safety and compliance capability, not documentation overhead.
Assignee: Unassigned · due 203 days after project start · takes 7 days
- #18Monitoring, Observability, and Drift DetectionMediumTo Do
Sets the maturity bar for detecting distribution shift, model staleness, and production degradation across your data and AI estate. You learn to build drift detection into operations rather than bolt it on.
Assignee: Unassigned · due 210 days after project start · takes 7 days
- #19Scalability and PerformanceMediumTo Do
Address how to hold system performance as load and use cases grow, with specific attention to protecting the safety and monitoring properties that held at small volume. You get the moves that keep scale from silently eroding guarantees.
Assignee: Unassigned · due 217 days after project start · takes 7 days
- #20Model Output / Prediction QualityHighTo Do
Define what 'good enough' output means for safety-critical use and how you orchestrate the levers — data, prompts, retrieval, adaptation — that move it across the portfolio.
Assignee: Unassigned · due 224 days after project start · takes 14 days
- #21System Reliability and MaintainabilityHighTo Do
Operational reliability as inseparable from safety and shows how to embed incident-readiness into how teams run AI systems under load and fault.
Assignee: Unassigned · due 238 days after project start · takes 14 days
- #22Evaluation Rigor and ReliabilityHighTo Do
Define what counts as a valid, reliable pre-deployment evaluation and how you defend it as the non-negotiable gate everything ships through.
Assignee: Unassigned · due 252 days after project start · takes 14 days
- #23Responsible AI and Fairness PracticesHighTo Do
Translating fairness, privacy, and harm-mitigation principles into enforceable practice that teams internalize during design — not compliance artifacts bolted on before launch.
Assignee: Unassigned · due 266 days after project start · takes 14 days
- #24Human-AI Collaboration ModeHighTo Do
Focus on the design guide for how humans and AI split work — the oversight roles, centaur and cyborg modes, and the collaboration patterns that keep human judgment load-bearing on high-stakes tasks.
Assignee: Unassigned · due 280 days after project start · takes 14 days
- #25Human Trust in AIHighTo Do
Focus on about calibrating trust — building enough confidence in AI to use it, without the blind faith that lets errors through or the reflexive rejection that wastes it.
Assignee: Unassigned · due 294 days after project start · takes 14 days
- #26AI Over-RelianceHighTo Do
Names uncritical dependence on AI as a measurable governance risk and shows how to design oversight that keeps human judgment from quietly eroding.
Assignee: Unassigned · due 308 days after project start · takes 14 days
- #27Systematic Developer IterationMediumTo Do
Use the discipline to turn safety improvement into a repeatable loop instead of a series of one-off heroic fixes. You learn how to make build-measure-learn the default cadence across teams.
Assignee: Unassigned · due 322 days after project start · takes 7 days
- #28Rigorous Investigation and Evidence DesignHighTo Do
Use the standard for evidence that actually demonstrates a safety property rather than gesturing at one.
Assignee: Unassigned · due 329 days after project start · takes 14 days
- #29Deliberate / Effortful ReasoningMediumTo Do
The effortful override — the slow, explicit checking you apply when a fast read isn't good enough to bet on.
Assignee: Unassigned · due 343 days after project start · takes 7 days
- #30Cognitive BiasesMediumTo Do
Names the systematic errors — anchoring, availability, base-rate neglect — that quietly warp organizational risk assessment, and how to blunt them.
Assignee: Unassigned · due 350 days after project start · takes 7 days
- #31Decision Effectiveness / Sound JudgmentMediumTo Do
Focus on about the end product all the judgment machinery serves: well-timed, well-reasoned go/no-go calls under real uncertainty.
Assignee: Unassigned · due 357 days after project start · takes 7 days
- #32Business Value and User SatisfactionHighTo Do
Keeps you honest that safety exists to enable useful systems, and shows how to hold value and safety together rather than trading them off.
Assignee: Unassigned · due 364 days after project start · takes 14 days
- #33Guardrails and Responsible AI GovernanceHighTo Do
Define the policy layer above technical controls — the responsible-AI governance framework that decides what your org will and won't deploy, and the leadership posture that enforces it.
Assignee: Unassigned · due 378 days after project start · takes 14 days
- #34Jagged Frontier KnowledgeHighTo Do
Build and share an org-wide map of where AI is surprisingly capable versus surprisingly brittle — the knowledge that governs what you dare to delegate.
Assignee: Unassigned · due 392 days after project start · takes 14 days
- #35Societal Impact and RiskHighTo Do
Keeps the population-level, long-horizon view of harms, fairness, and information integrity in scope even when your local product metrics look healthy.
Assignee: Unassigned · due 406 days after project start · takes 14 days
- #36Leadership Buy-In and Strategic AlignmentHighTo Do
Secure real executive sponsorship for safety by translating it into strategic terms leadership can act on and defend.
Assignee: Unassigned · due 420 days after project start · takes 14 days
- #37Workforce and Organizational AdaptationMediumTo Do
Make safe AI practice a property of the organization rather than a hobby of a few motivated individuals. You get the moves for reshaping roles, skills, and workflows so governance survives turnover.
Assignee: Unassigned · due 434 days after project start · takes 7 days
- #38Founder Vision and Conviction LeadershipHighTo Do
Focus on about setting a credible, ambitious vision for safety and alignment and holding conviction when defending it becomes costly. You get how vision provides direction and the resolve to keep it under pressure.
Assignee: Unassigned · due 441 days after project start · takes 14 days
- #39Regulatory StrategyHighTo Do
Treats the path through AI regulation as something you actively engineer — mapping pathways, structuring evidence, and shaping regulator engagement. You learn to use regulatory strategy as a design lever from the start.
Assignee: Unassigned · due 455 days after project start · takes 14 days
- #40Regulatory Approval and Market AuthorizationMediumTo Do
Attaining and maintaining authorization to deploy AI in regulated contexts. You learn why approval is a stateful commitment you keep re-earning, not a one-time gate you clear.
Assignee: Unassigned · due 469 days after project start · takes 7 days
- #41Integrated Multidisciplinary TeamsMediumTo Do
Compose and empower a safety team that runs alignment, engineering, and governance in parallel rather than handing work down a chain.
Assignee: Unassigned · due 476 days after project start · takes 7 days
- #42Intuitive / Recognitional JudgmentHighTo Do
Focus on about the trained read — sensing the shape of risk in a system or incident before you can fully explain it — and, crucially, knowing when to trust it.
Assignee: Unassigned · due 483 days after project start · takes 14 days
- #43Domain Experience BaseMediumTo Do
The accumulated store of incidents, near-misses, and analogues that feeds sound intuition — and how to grow it deliberately across the practice, not just in your own head.
Assignee: Unassigned · due 497 days after project start · takes 7 days