Metrics and Reporting
This document defines the metrics and reporting model for the AI Control Architecture.
The purpose is to help organizations measure AI control coverage, control maturity, evidence completeness, assurance results, exceptions, incidents, and improvement over time.
AI control reporting should not measure AI adoption alone.
High AI usage does not mean high AI control maturity.
The enterprise should measure whether AI is:
Visible
Owned
Risk-tiered
Controlled
Tested
Evidenced
Monitored
Containable
Improving
1. Purpose
The metrics and reporting model helps answer:
What AI exists?
Who owns it?
Which AI use cases are high risk?
Which controls are implemented?
Which controls are missing?
Which AI use cases have been tested?
Which evidence exists?
Which exceptions are open?
Which incidents occurred?
Which vendors create evidence gaps?
Where should leadership invest next?
Metrics should support action.
A metric that does not drive ownership, remediation, prioritization, or governance decision-making should be reviewed.
2. Reporting Principles
Principle 1: Measure Control, Not Just Adoption
Do not report only:
Number of AI users
Number of prompts
Number of AI tools deployed
Number of copilots enabled
Also report:
Number of AI use cases inventoried
Percentage with owners
Percentage risk-tiered
Percentage with required controls
Percentage with assurance completed
Percentage with evidence package
Percentage with containment path
Principle 2: Report by Risk Tier
Metrics should distinguish low-risk AI from high-risk AI.
A missing control on a Tier 1 drafting assistant is not the same as a missing control on a Tier 5 autonomous or regulated AI use case.
Report separately for:
Tier 1: Low-risk productivity or public-data use
Tier 2: Internal productivity with enterprise data
Tier 3: Decision-supporting AI
Tier 4: Action-capable AI
Tier 5: High-impact autonomous or regulated AI
Principle 3: Report by Pattern
Different AI patterns create different control issues.
Report metrics by AI pattern where useful:
Copilot
RAG system
Internal LLM application
AI-enabled SaaS
Embedded vendor AI
Agent
Customer-facing AI
Developer AI tool
Security operations AI
Decision-supporting AI
Action-capable AI
Principle 4: Separate Design, Operation, and Evidence
A control may be:
Designed
Implemented
Operating
Tested
Evidenced
Effective
Do not treat these as the same.
For example, a logging requirement may be designed but not yet producing useful evidence.
Principle 5: Metrics Must Have Owners
Every reported metric should have:
Metric owner
Data source
Reporting frequency
Target
Threshold
Action when outside threshold
3. Reporting Audiences
Different audiences need different reports.
4. Executive Dashboard
The executive dashboard should be short and risk-focused.
Recommended Executive Metrics
Executive Dashboard Format
AI Control Executive Dashboard
Reporting period:
Prepared by:
Overall status:
1. AI visibility
2. High-risk AI exposure
3. Control readiness
4. Assurance and evidence
5. Exceptions and findings
6. Incidents and near misses
7. Vendor AI risk
8. Maturity trend
9. Key decisions required
10. Investment or resource needs
5. Governance Dashboard
The governance dashboard should support operational decision-making.
Recommended Governance Metrics
Governance Reporting Questions
The governance report should answer:
Which AI use cases need decisions?
Which high-risk use cases are blocked?
Which exceptions are overdue?
Which findings need escalation?
Which vendors create unresolved risk?
Which AI use cases should be suspended, restricted, or retired?
6. Inventory Metrics
Inventory metrics measure AI visibility.
Core Inventory Metrics
Inventory Breakdown
Report inventory by:
Business unit
Function
Owner
AI pattern
Risk tier
Lifecycle stage
Vendor involvement
Customer-facing status
Action capability
Decision impact
Inventory Warning Indicators
Investigate when:
AI use cases have no owner.
AI use cases have unknown lifecycle status.
AI-enabled SaaS is not recorded.
Vendor AI features are discovered outside review.
High-risk AI is missing from the inventory.
7. Risk Tier Metrics
Risk tier metrics measure whether AI use cases are classified consistently.
Core Risk Metrics
Risk Distribution Report
Tier 1:
Tier 2:
Tier 3:
Tier 4:
Tier 5:
Unknown:
Risk Tier Warning Indicators
Investigate when:
A use case has unknown risk tier.
A customer-facing use case is Tier 1 or Tier 2 without rationale.
An action-capable use case is below Tier 4.
A regulated or high-impact use case is below Tier 5.
Risk tier has not been reviewed after major change.
8. Control Coverage Metrics
Control coverage metrics measure whether required controls are defined and implemented.
Pillar Coverage Metrics
Control Coverage by Tier
Report control coverage by risk tier.
Tier 1 control coverage:
Tier 2 control coverage:
Tier 3 control coverage:
Tier 4 control coverage:
Tier 5 control coverage:
Control Coverage Warning Indicators
Investigate when:
Tier 3 AI lacks decision owner.
Tier 4 AI lacks tool/action control.
Tier 5 AI lacks evidence package.
Action-capable AI lacks kill switch.
Vendor AI lacks vendor assessment.
Sensitive data AI lacks data boundary.
9. Identity and Access Metrics
Identity metrics measure AI authority control.
Core Identity Metrics
Identity Warning Indicators
Investigate when:
AI access uses shared accounts.
AI service accounts are over-permissioned.
AI has privileged access without approval.
Agent identity cannot be disabled quickly.
Vendor AI identity model is unknown.
AI activity cannot be distinguished from human activity.
10. Data Boundary Metrics
Data metrics measure whether AI data access is controlled.
Core Data Metrics
RAG-Specific Data Metrics
Data Warning Indicators
Investigate when:
Data classification is unknown.
Data owner approval is missing.
Sensitive data is used without boundary.
RAG system lacks retrieval testing.
Vendor retention is unknown.
Training/reuse setting is unknown.
11. Prompt and Input Metrics
Prompt and input metrics measure whether input risk is managed.
Core Prompt/Input Metrics
Prompt/Input Warning Indicators
Investigate when:
System prompts are not versioned.
Prompt changes bypass review.
External content is processed without prompt injection testing.
Sensitive data inputs are not detected or blocked.
Tool responses are treated as trusted instructions.
12. Output and Decision Metrics
Output and decision metrics measure whether AI output is controlled before it becomes a decision, record, communication, or action.
Core Output Metrics
Decision Metrics
Output/Decision Warning Indicators
Investigate when:
AI recommendation becomes final decision.
Decision owner is missing.
Human review has near-zero rejection or modification rate.
Generated records lack provenance.
Customer-facing outputs lack correction path.
13. Tool and Action Metrics
Tool and action metrics measure whether AI can affect enterprise state safely.
Core Tool/Action Metrics
Agent-Specific Metrics
Tool/Action Warning Indicators
Investigate when:
Tool inventory is incomplete.
AI can perform actions without approval.
High-risk actions lack logs.
Kill switch is missing or untested.
Rollback is unknown.
Agent retries or tool calls spike unexpectedly.
14. Vendor AI Metrics
Vendor AI metrics measure third-party AI risk.
Core Vendor Metrics
Vendor Warning Indicators
Investigate when:
Vendor AI feature is enabled by default.
Training/reuse setting is unknown.
Vendor retains prompts or outputs without approval.
Logs are unavailable.
Vendor incident support path is unclear.
Vendor changed AI feature behavior without review.
15. Assurance Metrics
Assurance metrics measure whether AI behavior and controls have been tested.
Core Assurance Metrics
Assurance Warning Indicators
Investigate when:
Tier 4 or Tier 5 AI has no assurance test plan.
Prompt injection testing is missing for RAG or agentic AI.
Tool/action testing is missing for action-capable AI.
Evidence reconstruction has not been tested.
Findings remain overdue.
Regression testing is not triggered after material change.
16. Evidence Metrics
Evidence metrics measure auditability and incident readiness.
Core Evidence Metrics
Evidence Warning Indicators
Investigate when:
High-risk AI lacks evidence package.
AI activity cannot be reconstructed.
Approval records are missing.
Vendor logs are unavailable.
Evidence retention is shorter than business, legal, or incident needs.
17. Exception Metrics
Exception metrics measure accepted control gaps.
Core Exception Metrics
Exception Warning Indicators
Investigate when:
Exception has no expiry date.
Exception has no owner.
Exception lacks compensating control.
Exception is repeatedly renewed.
Exception relates to Tier 4 or Tier 5 AI.
Exception contributed to incident.
18. Incident Metrics
Incident metrics measure AI operational risk and resilience.
Core Incident Metrics
Incident Warning Indicators
Investigate when:
Incident cannot be reconstructed.
Containment is delayed.
Kill switch fails.
Vendor evidence is unavailable.
Same failure pattern repeats.
Post-incident actions are not closed.
19. Maturity Metrics
Maturity metrics show progress across the architecture.
Pillar Maturity Metrics
Report current and target maturity for each pillar:
Maturity Warning Indicators
Investigate when:
High-risk adoption grows faster than control maturity.
Tool/action capability grows while incident containment remains low.
Vendor AI use grows while vendor evidence remains weak.
Decision-supporting AI grows while accountability remains low.
AI incidents increase without control improvements.
20. Reporting Cadence
A suggested reporting cadence is below.
21. Sample Monthly Governance Report
Use this structure for a monthly governance report.
AI Control Governance Report
Reporting period:
Prepared by:
Overall status:
1. Executive summary
2. New AI use cases
3. AI inventory status
4. Risk tier distribution
5. High-risk AI review status
6. Control coverage
7. Vendor AI status
8. Assurance testing status
9. Evidence package status
10. Open findings
11. Open exceptions
12. Incidents and near misses
13. Maturity changes
14. Decisions required
15. Escalations
16. Next-period priorities
22. Sample Executive Summary
Example:
During this reporting period, the enterprise inventory increased from [X] to [Y] AI use cases. [Z]% of use cases now have assigned business owners, and [Z]% have assigned risk tiers.
There are currently [X] Tier 3, [Y] Tier 4, and [Z] Tier 5 AI use cases. [N] high-risk use cases still require control assessment.
The most significant control gaps are [gap 1], [gap 2], and [gap 3]. There are [N] open high-severity findings and [N] open exceptions, of which [N] are overdue.
There were [N] AI incidents and [N] near misses this period. The most important incident theme was [theme].
The recommended leadership actions are:
1. [Action]
2. [Action]
3. [Action]
23. Metric Data Quality
Metrics are only useful if the data is reliable.
For each metric, define:
Metric name:
Metric owner:
Data source:
Calculation method:
Reporting frequency:
Target:
Threshold:
Known limitations:
Action when metric is outside threshold:
Data Quality Warning Indicators
Investigate when:
Metric source is manual and inconsistent.
Metric owner is unclear.
Definitions vary across teams.
Risk tier is not applied consistently.
Vendor evidence is self-reported but not validated.
Dashboard shows green despite incidents or exceptions.
24. Targets and Thresholds
Targets should be risk-tiered.
Example targets:
Targets should be adapted to the organization’s maturity stage.
25. Using Metrics to Drive Action
Metrics should trigger action.
Examples:
26. Reporting Anti-Patterns
Anti-Pattern 1: Adoption Theater
Reporting only adoption numbers.
Example:
10,000 users enabled for AI
1 million prompts submitted
Why this fails:
It does not show whether AI is controlled.
Anti-Pattern 2: Green Dashboard With No Evidence
Showing controls as complete because a document exists.
Why this fails:
A control is not effective until it is implemented, tested, and evidenced.
Anti-Pattern 3: No Risk Tier Distinction
Combining low-risk drafting tools and high-risk autonomous agents in the same metric.
Why this fails:
It hides serious risk concentration.
Anti-Pattern 4: Vendor Blind Spots
Reporting vendor AI as approved without tracking logs, retention, training/reuse, and incident support.
Why this fails:
Vendor AI can create enterprise risk even when internally built AI is controlled.
Anti-Pattern 5: Ignoring Near Misses
Tracking only confirmed incidents.
Why this fails:
Near misses reveal control weakness before major harm occurs.
27. Summary
AI control reporting should show whether the enterprise can safely adopt AI.
The best metrics answer:
Do we know what AI exists?
Is it owned?
Is it risk-tiered?
Are required controls implemented?
Are controls tested?
Does evidence exist?
Are exceptions temporary?
Are incidents contained?
Are lessons improving the architecture?
The goal of reporting is not to create dashboards.
The goal is to create decisions, accountability, remediation, and safer AI adoption.