← Templates

AI Safety, Governance & Alignment Practice (FIELD GUIDE (ALL LEVELS))

Start this plan

The central quality signal the field guide holds teams accountable to — defining what 'good enough' output means for safety-critical use and orchestrating the levers (data, prompts, retrieval, adaptation) that move it across the portfolio. (Setting the org's safety standard and defending it as a field guide.)

started 0 · finished 0 (claimed) · not yet measured (verified) · no data (n<5)

Derived from: AI Safety, Governance & Alignment Practice (FIELD GUIDE (ALL LEVELS))T1

Ordered tasks (43) — this is what a project auto-creates

  • #1Model and Provider SelectionHighTo Do

    Use the criteria and escalation frame for choosing models, providers, and architectures across a portfolio of safety-critical systems — not for a single build.

    Assignee: Unassigned · due 0 days after project start · takes 14 days

  • #2Training Data Quality, Coverage, and QuantityHighTo Do

    Define the data-quality standard you hold across teams — treating provenance, coverage, and bias exposure as governance obligations rather than engineering nice-to-haves.

    Assignee: Unassigned · due 14 days after project start · takes 14 days

  • #3Hallucination RateHighTo Do

    Turn 'the model makes things up' into a governed metric with risk-tiered thresholds and known reduction levers.

    Assignee: Unassigned · due 28 days after project start · takes 14 days

  • #4Prompt Engineering QualityHighTo Do

    Use the practices for making prompt design a reviewed, versioned discipline rather than individual craft that varies by author.

    Assignee: Unassigned · due 42 days after project start · takes 14 days

  • #5Data Warehouse / System Architecture DesignHighTo Do

    Set the structural foundation — storage, partitioning, replication, and access patterns — beneath safety-critical models. You get the case for making architecture an org-wide standard rather than a per-project choice.

    Assignee: Unassigned · due 56 days after project start · takes 14 days

  • #6Data Integration and ETL QualityHighTo Do

    Address the rigor of the pipelines that combine heterogeneous sources into the data your models consume. You learn why pipeline integrity is a precondition for any credible safety claim.

    Assignee: Unassigned · due 70 days after project start · takes 14 days

  • #7Data Quality and Information TrustHighTo Do

    Owning the accuracy, completeness, currency, and conformance of your delivered data — and the confidence stakeholders reasonably place in it. You get the case for treating trust as a governed, measured property.

    Assignee: Unassigned · due 84 days after project start · takes 14 days

  • #8Requirements and Business Objective AlignmentMediumTo Do

    Anchor safety work to genuine problems and purposes rather than to whatever the available data and tooling make convenient. You get the practice of defining the problem before designing the solution.

    Assignee: Unassigned · due 98 days after project start · takes 7 days

  • #9Fine-Tuning and Model AdaptationHighTo Do

    Judge when adapting a pretrained model genuinely improves safety and when it silently degrades alignment already baked into the base model.

    Assignee: Unassigned · due 105 days after project start · takes 14 days

  • #10Observability and Monitoring MaturityHighTo Do

    Define what 'observable enough to be safe' means and how you ensure teams instrument systems to that bar before shipping, not after an incident.

    Assignee: Unassigned · due 119 days after project start · takes 14 days

  • #11Generalization / Model Variance ControlHighTo Do

    Equips you to reason about and communicate out-of-distribution robustness — distinguishing models that genuinely generalize from those that only look strong on the test set.

    Assignee: Unassigned · due 133 days after project start · takes 14 days

  • #12Deployment and Serving ArchitectureMediumTo Do

    Treats deployment and serving patterns as a safety boundary — architecting rollout gates, rollback paths, and serving robustness so unsafe behavior is contained before it reaches users.

    Assignee: Unassigned · due 147 days after project start · takes 7 days

  • #13Inference Efficiency and CostMediumTo Do

    Hold the line where cost and latency pressure meets safety margin — deciding when efficiency gains must not be allowed to thin out monitoring, guardrails, or evaluation.

    Assignee: Unassigned · due 154 days after project start · takes 7 days

  • #14Guardrails and AI Security ControlsHighTo Do

    Define the guardrail and security-control regime that constrains unsafe model behavior in production, and how you get it implemented consistently instead of team-by-team improvisation.

    Assignee: Unassigned · due 161 days after project start · takes 14 days

  • #15Data Distribution Shift and Model FreshnessHighTo Do

    Drift and staleness as an ongoing threat you must detect and respond to on a horizon longer than any single release.

    Assignee: Unassigned · due 175 days after project start · takes 14 days

  • #16Context Grounding / Retrieval RelevanceHighTo Do

    Sets the standard for how tightly outputs must be tied to retrieved evidence before they can be trusted in safety-critical use.

    Assignee: Unassigned · due 189 days after project start · takes 14 days

  • #17Metadata Management and Data GovernanceMediumTo Do

    Establish lineage, metadata, and governance so data provenance is traceable and auditable across the org. You learn why this is a foundational safety and compliance capability, not documentation overhead.

    Assignee: Unassigned · due 203 days after project start · takes 7 days

  • #18Monitoring, Observability, and Drift DetectionMediumTo Do

    Sets the maturity bar for detecting distribution shift, model staleness, and production degradation across your data and AI estate. You learn to build drift detection into operations rather than bolt it on.

    Assignee: Unassigned · due 210 days after project start · takes 7 days

  • #19Scalability and PerformanceMediumTo Do

    Address how to hold system performance as load and use cases grow, with specific attention to protecting the safety and monitoring properties that held at small volume. You get the moves that keep scale from silently eroding guarantees.

    Assignee: Unassigned · due 217 days after project start · takes 7 days

  • #20Model Output / Prediction QualityHighTo Do

    Define what 'good enough' output means for safety-critical use and how you orchestrate the levers — data, prompts, retrieval, adaptation — that move it across the portfolio.

    Assignee: Unassigned · due 224 days after project start · takes 14 days

  • #21System Reliability and MaintainabilityHighTo Do

    Operational reliability as inseparable from safety and shows how to embed incident-readiness into how teams run AI systems under load and fault.

    Assignee: Unassigned · due 238 days after project start · takes 14 days

  • #22Evaluation Rigor and ReliabilityHighTo Do

    Define what counts as a valid, reliable pre-deployment evaluation and how you defend it as the non-negotiable gate everything ships through.

    Assignee: Unassigned · due 252 days after project start · takes 14 days

  • #23Responsible AI and Fairness PracticesHighTo Do

    Translating fairness, privacy, and harm-mitigation principles into enforceable practice that teams internalize during design — not compliance artifacts bolted on before launch.

    Assignee: Unassigned · due 266 days after project start · takes 14 days

  • #24Human-AI Collaboration ModeHighTo Do

    Focus on the design guide for how humans and AI split work — the oversight roles, centaur and cyborg modes, and the collaboration patterns that keep human judgment load-bearing on high-stakes tasks.

    Assignee: Unassigned · due 280 days after project start · takes 14 days

  • #25Human Trust in AIHighTo Do

    Focus on about calibrating trust — building enough confidence in AI to use it, without the blind faith that lets errors through or the reflexive rejection that wastes it.

    Assignee: Unassigned · due 294 days after project start · takes 14 days

  • #26AI Over-RelianceHighTo Do

    Names uncritical dependence on AI as a measurable governance risk and shows how to design oversight that keeps human judgment from quietly eroding.

    Assignee: Unassigned · due 308 days after project start · takes 14 days

  • #27Systematic Developer IterationMediumTo Do

    Use the discipline to turn safety improvement into a repeatable loop instead of a series of one-off heroic fixes. You learn how to make build-measure-learn the default cadence across teams.

    Assignee: Unassigned · due 322 days after project start · takes 7 days

  • #28Rigorous Investigation and Evidence DesignHighTo Do

    Use the standard for evidence that actually demonstrates a safety property rather than gesturing at one.

    Assignee: Unassigned · due 329 days after project start · takes 14 days

  • #29Deliberate / Effortful ReasoningMediumTo Do

    The effortful override — the slow, explicit checking you apply when a fast read isn't good enough to bet on.

    Assignee: Unassigned · due 343 days after project start · takes 7 days

  • #30Cognitive BiasesMediumTo Do

    Names the systematic errors — anchoring, availability, base-rate neglect — that quietly warp organizational risk assessment, and how to blunt them.

    Assignee: Unassigned · due 350 days after project start · takes 7 days

  • #31Decision Effectiveness / Sound JudgmentMediumTo Do

    Focus on about the end product all the judgment machinery serves: well-timed, well-reasoned go/no-go calls under real uncertainty.

    Assignee: Unassigned · due 357 days after project start · takes 7 days

  • #32Business Value and User SatisfactionHighTo Do

    Keeps you honest that safety exists to enable useful systems, and shows how to hold value and safety together rather than trading them off.

    Assignee: Unassigned · due 364 days after project start · takes 14 days

  • #33Guardrails and Responsible AI GovernanceHighTo Do

    Define the policy layer above technical controls — the responsible-AI governance framework that decides what your org will and won't deploy, and the leadership posture that enforces it.

    Assignee: Unassigned · due 378 days after project start · takes 14 days

  • #34Jagged Frontier KnowledgeHighTo Do

    Build and share an org-wide map of where AI is surprisingly capable versus surprisingly brittle — the knowledge that governs what you dare to delegate.

    Assignee: Unassigned · due 392 days after project start · takes 14 days

  • #35Societal Impact and RiskHighTo Do

    Keeps the population-level, long-horizon view of harms, fairness, and information integrity in scope even when your local product metrics look healthy.

    Assignee: Unassigned · due 406 days after project start · takes 14 days

  • #36Leadership Buy-In and Strategic AlignmentHighTo Do

    Secure real executive sponsorship for safety by translating it into strategic terms leadership can act on and defend.

    Assignee: Unassigned · due 420 days after project start · takes 14 days

  • #37Workforce and Organizational AdaptationMediumTo Do

    Make safe AI practice a property of the organization rather than a hobby of a few motivated individuals. You get the moves for reshaping roles, skills, and workflows so governance survives turnover.

    Assignee: Unassigned · due 434 days after project start · takes 7 days

  • #38Founder Vision and Conviction LeadershipHighTo Do

    Focus on about setting a credible, ambitious vision for safety and alignment and holding conviction when defending it becomes costly. You get how vision provides direction and the resolve to keep it under pressure.

    Assignee: Unassigned · due 441 days after project start · takes 14 days

  • #39Regulatory StrategyHighTo Do

    Treats the path through AI regulation as something you actively engineer — mapping pathways, structuring evidence, and shaping regulator engagement. You learn to use regulatory strategy as a design lever from the start.

    Assignee: Unassigned · due 455 days after project start · takes 14 days

  • #40Regulatory Approval and Market AuthorizationMediumTo Do

    Attaining and maintaining authorization to deploy AI in regulated contexts. You learn why approval is a stateful commitment you keep re-earning, not a one-time gate you clear.

    Assignee: Unassigned · due 469 days after project start · takes 7 days

  • #41Integrated Multidisciplinary TeamsMediumTo Do

    Compose and empower a safety team that runs alignment, engineering, and governance in parallel rather than handing work down a chain.

    Assignee: Unassigned · due 476 days after project start · takes 7 days

  • #42Intuitive / Recognitional JudgmentHighTo Do

    Focus on about the trained read — sensing the shape of risk in a system or incident before you can fully explain it — and, crucially, knowing when to trust it.

    Assignee: Unassigned · due 483 days after project start · takes 14 days

  • #43Domain Experience BaseMediumTo Do

    The accumulated store of incidents, near-misses, and analogues that feeds sound intuition — and how to grow it deliberately across the practice, not just in your own head.

    Assignee: Unassigned · due 497 days after project start · takes 7 days