Back to Resources

Printable pilot resource

AI Police Report Pilot Scorecard

Define the baseline, record quality and workflow evidence, track failures and remediation, and make a written go, revise, or stop decision.

Set the rules before the first pilot case

Define the comparison baseline, unit of analysis, reviewer instructions, severity scale, minimum sample, decision thresholds, and stop conditions in writing. Use representative approved scenarios and the same review standard for the current workflow and the pilot workflow. This scorecard supports evaluation; it does not determine legal sufficiency, CJIS compliance, or whether a product is appropriate for an agency.

Read the pilot planning guide
A

Pilot record and comparison design

Complete one cover sheet for the pilot and identify any changes to scope, reviewers, or measurement rules as they occur.

Agency / unit
Pilot owner
Pilot period
Version / configuration
In-scope report types
Scenario and operational context
Baseline workflow and date range
Participating officers / roles
Independent reviewers / roles
Minimum sample and exclusions
Severity definitions and adjudication process
Prewritten success thresholds
Immediate pause / stop conditions
01

Time and workflow effort

Use the same start and end points for the current process and the pilot. Generation speed alone does not measure the work required to produce a supervisor-ready draft.

End-to-end completion time

Elapsed time from the agreed starting event to a supervisor-ready draft, including waiting, review, corrections, and export steps.

Record as: Median minutes; also note the range and outliers

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location

Active editing time

Time the officer actively spends reviewing, correcting, adding, deleting, or reorganizing draft language before it is ready for supervisor review.

Record as: Median minutes and percentage of total time

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location

Supervisor returns

Drafts returned for factual, clarity, completeness, attribution, policy, formatting, or other corrections. Record the reason for every return.

Record as: Count and percentage of reviewed drafts, by reason

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location
02

Factual integrity and completeness

Have trained reviewers compare each pilot draft with the approved source set and the agency's normal review standard. Count findings consistently and retain examples for adjudication.

Unsupported statements

Statements in the draft that are not supported by the officer-provided facts or another approved, identified source. Do not combine these with style preferences.

Record as: Findings per draft and percentage of drafts affected

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location

Material omissions

Relevant supplied facts absent from the draft when the agency's review standard calls for their inclusion. Record why the omission matters to documentation quality.

Record as: Findings per draft, severity, and percentage affected

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location

Attribution or source errors

Language assigned to the wrong person, event, document, policy, instruction, or other source, including an incorrect or untraceable source reference.

Record as: Count by error type, severity, and source category

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location
03

Reliability and user experience

Separate system reliability from output quality, and collect structured feedback alongside open comments. A favorable comment does not offset an unresolved material finding.

System failures

Attempts that could not complete because of a product, model-provider, identity, network, export, or integration issue. Record duration, impact, recovery, and recurrence.

Record as: Failure rate, unavailable minutes, and affected workflows

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location

User feedback

Officer and reviewer ratings for usefulness, clarity of prompts, review burden, trust calibration, training sufficiency, and workflow fit, plus categorized comments.

Record as: Rating distribution, response count, and recurring themes

Baseline
Pilot result
Threshold
Status
Met Concern
Notes / evidence location
B

Finding and remediation log

Record each finding separately. Include enough context for a reviewer who did not observe the original workflow to reproduce or adjudicate it.

Finding types

Unsupported statement · Material omission · Attribution/source error · Supervisor return · System failure · User-reported issue

Finding 1

ID: __________
Date / case or scenario ID
Finding type
Scenario and operational context
Source set and version
Reviewer and role
Finding description and evidence
Severity and rationale
Immediate disposition
Remediation, owner, and due date
Retest result / closure evidence

Finding 2

ID: __________
Date / case or scenario ID
Finding type
Scenario and operational context
Source set and version
Reviewer and role
Finding description and evidence
Severity and rationale
Immediate disposition
Remediation, owner, and due date
Retest result / closure evidence

Finding 3

ID: __________
Date / case or scenario ID
Finding type
Scenario and operational context
Source set and version
Reviewer and role
Finding description and evidence
Severity and rationale
Immediate disposition
Remediation, owner, and due date
Retest result / closure evidence

Written decision gate

Go, revise and re-test, or stop

Compare the results with the thresholds approved before the pilot. Document exceptions and unresolved findings; do not treat an overall average as a substitute for reviewing high-severity failures.

Go

Approved scope may proceed subject to the conditions below.

Revise and re-test

Remediation and a defined re-test are required before expansion.

Stop

The pilot will not proceed in its current form.

Decision date
Scope covered by this decision
Written rationale, including unmet thresholds
Unresolved findings and accepted residual risk
Required remediation and re-test plan
Conditions, owners, and due dates
Operational owner approval
Quality / reviewer approval
Security / privacy review
Agency-authorized decision maker

Keep the evaluation boundary explicit

Results apply only to the tested scenarios, users, product version, configuration, sources, deployment, and review process. Reassess material changes and validate agency requirements with the appropriate operational, security, privacy, procurement, and legal reviewers.

Planning references

LeoPen resource · Updated September 1, 2026