Owen Sound: A Four-Year City Business Plan

Home › Appendices › Appendix G

Appendices

Appendix GScorecard Definitions and Measurement Standards

10,737 words · Mike Seiler · Owen Sound, Ontario

Open in the reader →

Vote on the proposals, hear the audio, read the reviews, search the whole plan.

In this chapter

One definition. One baseline. One method. One public record.

A public Scorecard is useful only if the numbers mean the same thing:

Without fixed definitions, government can make almost any record look successful.

It can:

This appendix exists to prevent that.

The governing principle is:

Define the measure before judging the result.

The public standard should be:

One definition. One baseline. One method. One public record.

If a definition must change:

change it openly, explain why, preserve the old result and restate prior periods where practical.

The Scorecard should never become:

It should answer:

What happened?

Compared with what?

How do we know?

How confident are we?

What remains unknown?

What should happen next?

G.1Purpose

This appendix establishes the common measurement language behind Sections 44 through 60.

G.2Scorecard Sections

The public Scorecard covers:

Finance

Tax

Services

Infrastructure

Downtown

Housing

Business

Safety

Recreation

Youth

Accessibility

Privacy and Information

Environment

Partnership

Resident Satisfaction

Completed Commitments

Missed Commitments

G.3One Measurement Standard

Departments may use specialized operational measures.

Public headline measures should still follow:

G.4Technical Measures Can Remain Technical

A bridge engineer may need:

A public Scorecard may summarize them.

It must not:

G.5Measurement Is Not Management by Spreadsheet

Numbers inform:

They do not replace:

G.6Metric

A metric is:

a defined measure used to describe an activity, output, outcome, condition, cost, risk or experience.

G.7Measure

The terms:

may generally be used interchangeably in this plan.

G.8Headline Measure

A limited number of measures selected for:

G.9Supporting Measure

Provides:

G.10Diagnostic Measure

Helps explain:

why a headline measure moved.

G.11No Metric Explosion

Do not publish hundreds of numbers merely because:

G.12Measure What Matters

A good Scorecard uses enough information to:

without burying residents.

G.13Metric Definition Card

Every recurring headline measure should have a:

Metric Definition Card.

G.14Required Metric Definition Fields

Each card should identify:

Metric Name

Plain-Language Question

Definition

Purpose

Unit

Numerator

Denominator where applicable

Inclusion Rules

Exclusion Rules

Data Source

Data Owner

Measurement Frequency

Reporting Frequency

Baseline

Target where applicable

Direction of Improvement

Confidence

Known Limitations

Revision History

G.15Example

Metric:

Median City-Controlled Development Review Time

Plain-language question:

How long is the City's part of the review taking?

Definition must state:

G.16Stable Definition

Once adopted:

Keep the definition stable unless there is:

to change it.

G.17Methodological Improvement

Allowed.

G.18Political Convenience

Not a methodological reason.

G.19Baseline

A baseline is:

the defined reference point against which later performance is compared.

G.20Baseline Date

Every baseline should identify:

G.21Baseline Can Be

G.22Baseline Choice

Should match:

G.23Seasonal Measure

May require:

G.24Winter Maintenance

Do not compare:

directly to:

without context.

G.25Downtown Pedestrian Activity

May require:

G.26Baseline Before Intervention

Preferably establish before:

G.27Missing Baseline

If no credible baseline exists:

Say:

Baseline unavailable. Year One establishes baseline.

G.28Do Not Reconstruct False Precision

Do not invent exact historical values from:

G.29Estimated Baseline

Can be used where useful.

Must be labelled:

Estimated.

G.30Revised Baseline

If original baseline materially wrong:

Preserve:

G.31No Quiet Rebase

Never.

G.32Target

A target is:

a defined desired result for a defined period.

G.33Target Is Not Forecast

Target says:

what we intend to achieve.

Forecast says:

what we currently expect to happen.

G.34Target Is Not Guarantee

No.

G.35Targets Should Be Defensible

Based on:

G.36Round-Number Target

May be reasonable.

But not merely because:

G.37No Target

Some measures should be monitored without:

G.38Example

Number of privacy incidents.

The desired direction may be:

But setting:

zero reports

could discourage reporting.

G.39Target Range

Sometimes better than:

G.40Minimum Standard

Some measures have:

minimums.

G.41Aspirational Target

May exceed:

Label.

G.42Actual

An actual is:

an observed or recorded result for a completed period or event.

G.43Actual Must Be Real

Do not call:

result:

G.44Preliminary Actual

Can be used before final reconciliation.

Label:

Preliminary.

G.45Final Actual

After appropriate:

G.46Forecast

A forecast is:

the best current estimate of a future result based on available information.

G.47Forecast Date

Show.

G.48Forecast Can Change

Yes.

G.49Forecast Revision

Not automatically:

G.50Forecast Accuracy

Can itself be measured where useful.

G.51Estimate

An estimate is:

an approximate value derived from incomplete information, modelling, sampling or professional judgment.

G.52Estimate Must Be Labelled

Always where material.

G.53Modelled

Use when result comes from:

G.54Survey Estimate

Use when extrapolated from:

G.55Administrative Record

Use when derived from:

G.56External Data

Use when data originates outside City.

G.57Source Matters

Every material Scorecard result should identify:

G.58Source Date

Where external:

Show latest available period.

G.59Stale External Data

Do not present as:

G.60Direction of Improvement

Every metric should state whether:

Higher Is Generally Better

Lower Is Generally Better

Target Range Is Better

Context Only

G.61Context-Only Measure

Important.

G.62Example

Population.

Higher is not automatically:

G.63Police Calls for Service

Higher is not automatically:

G.64Business Openings

Higher may be positive.

But requires:

G.65Number of Complaints

Higher could mean:

or:

G.66No Automatic Interpretation

Use context.

G.67Activity

An activity is:

something government or a partner does.

Examples:

G.68Output

An output is:

a direct product or deliverable of the activity.

Examples:

G.69Outcome

An outcome is:

the change or public result the activity is intended to produce.

Examples:

G.70Impact

A longer-term change that may involve:

G.71Attribution

The City should claim:

carefully.

G.72Caused

Use only where evidence reasonably supports:

G.73Contributed To

Often more accurate.

G.74Coincided With

Use when relationship is:

G.75Do Not Claim Outcome From Activity Alone

Meeting held does not prove:

G.76Do Not Claim Impact From Output Alone

Permit issued does not prove:

G.77Count

A count is:

the number of defined events, items or people meeting the stated criteria.

G.78Count Needs Definition

Count of:

G.79Unique Count

A count where duplicate entities are:

G.80Unique Resident

Use cautiously.

G.81Privacy

Do not build unnecessary persistent resident identity systems merely to:

G.82Privacy-Safe Approximation

May be preferable.

G.83Registration

Means:

G.84Attendance

Means:

G.85Completion

Means:

G.86Unique Participant

Means:

if privacy-safe methodology allows.

G.87Repeat Participation

Can be valuable.

Report separately where relevant.

G.88Do Not Call 1,000 Visits 1,000 People

Unless they actually are.

G.89Rate

A rate is:

a count relative to a defined population, exposure, time or opportunity.

G.90Rate Needs Denominator

Always.

G.91Denominator

The population or quantity against which numerator is measured.

G.92Numerator

The events or observations being counted.

G.93Denominator Drift

Dangerous.

G.94Example

Complaint rate can change because:

Show both where useful.

G.95Percentage

Percentage = Numerator ÷ Denominator × 100

G.96Percentage Point

Different from:

G.97Example

Moving from 40% to 50% is:

G.98Do Not Confuse Them

No.

G.99Ratio

Define:

G.100Per Capita

Define population source and:

G.101Population Lag

External population estimates may lag.

Label.

G.102Per Household

Define household source.

G.103Per Kilometre

Define:

correctly.

G.104Road Kilometres

Do not mix:

kilometres.

G.105Per Application

Define application type.

G.106Per Service Request

Define service request.

G.107Mean

Arithmetic average.

G.108Median

Middle value when observations are ordered.

G.109Median Is Often Useful for Service Time

Because a few extreme cases can distort:

G.110Mean Still Useful

Can show total workload effect.

G.111Use Both Where Helpful

G.11290th Percentile

The value at or below which approximately 90% of observations fall.

G.113Why Use It

To reveal:

hidden by median.

G.11495th Percentile

Can be used for highly operational services.

G.115Do Not Choose Percentile After Results

Define before.

G.116Minimum

Can hide everything.

G.117Maximum

Can be distorted by one exceptional case.

G.118Distribution

Sometimes more informative.

G.119Small Sample

Be careful with:

G.120Small-N Suppression

May also protect:

G.121Time

Time measures require:

G.122Calendar Day

Includes:

G.123Business Day

Must define which days count.

G.124Service Hour

May be more appropriate for some operational measures.

G.125Clock Start

Define.

G.126Clock Stop

Define.

G.127Clock Pause

Define.

G.128No Retroactive Clock Rules

Never.

G.129Resident Wait Time

Time experienced by:

G.130City-Controlled Time

Time while the file is actively within:

G.131Applicant Delay

Time waiting for:

G.132External Delay

Time waiting for:

where applicable.

G.133Total Elapsed Time

Can still matter to resident.

G.134Report Both Where Useful

Example:

Total elapsed

City-controlled

G.135Do Not Pause Clock Without Reason

No.

G.136Incomplete Application

Must have:

G.137"Complete"

Needs:

G.138Request Date

Not automatically same as:

G.139First Response Time

Must define what qualifies as:

response.

G.140Automated Receipt

Should not count as:

unless metric explicitly measures:

G.141Meaningful Response

Could mean:

G.142Closure

A service request is closed when:

the defined service process has reached a legitimate resolution or disposition.

G.143Closed Does Not Mean Resident Happy

No.

G.144Closed Does Not Mean Problem Fixed

Could be:

G.145Closure Code

Use.

G.146Example Closure Codes

Completed

No Issue Found

Duplicate

Referred

Outside Authority

Scheduled for Capital

Applicant Withdrew

Other Defined Disposition

G.147Closed for Statistics

Not allowed.

G.148Reopened

Track.

G.149Repeat Request

Track where appropriate.

G.150Repeat Issue

Different from:

G.151No Resident Blame Metric

Do not label residents:

for public performance analysis.

G.152Backlog

A backlog is:

open work that has passed the point at which it would normally or reasonably be expected to be completed under the defined service standard or work plan.

G.153Open Work Is Not Automatically Backlog

Important.

G.154Scheduled Future Work

May be:

not backlog.

G.155Backlog Definition

Should be service-specific.

G.156Backlog Count

Track.

G.157Backlog Age

Track.

G.158Oldest Backlog

Track where useful.

G.159High-Priority Backlog

Track separately.

G.160Backlog Dollar Value

Only where defensible.

G.161Do Not Clear Backlog by Deleting Items

No.

G.162Transfer

A service transfer is:

movement of a resident or file from one responsible service point to another.

G.163Internal Transfer

Within:

G.164External Handoff

To:

G.165Warm Handoff

Means more than:

call this number.

It may include:

where appropriate.

G.166Warm Handoff Does Not Mean Sharing Entire File

Privacy applies.

G.167First-Contact Resolution

A request resolved at first meaningful service contact without unnecessary transfer.

G.168Define "Resolved"

Service-specific.

G.169First-Contact Resolution Is Not Always Better

Complex matters may properly require:

G.170Wrong-Door Handoff

Resident initially contacts an institution not responsible for:

G.171Wrong-Door Rate

Could measure:

G.172Fewer Handoffs Is Not Automatically Better

If City simply rejects:

G.173Pair With

G.174Condition

Condition describes:

the physical or functional state of an asset according to a defined methodology.

G.175Condition Categories

Possible:

Very Good

Good

Fair

Poor

Very Poor

Unknown

G.176Define Each Category by Asset Class

Road:

G.177Do Not Invent Universal Condition Formula

No.

G.178Condition Source

Could be:

G.179Visual Observation

May support preliminary rating.

Label confidence.

G.180Confidence

Confidence describes:

how reliable or current the underlying information is.

G.181Confidence Categories

Possible:

High

Moderate

Low

Unknown

G.182High Confidence

Recent and appropriate:

G.183Moderate Confidence

Reasonably reliable but:

G.184Low Confidence

Limited or indirect evidence.

G.185Unknown

Insufficient basis.

G.186Confidence Is Not Condition

Critical.

G.187Good, Low Confidence

Should not be treated like:

G.188Unknown Is Not Failure

It is:

G.189Unknown Can Be Risk

Especially for:

G.190Criticality

The consequence of:

failure.

G.191Criticality Is Not Condition

Again.

G.192Risk

Risk generally combines:

according to the subject.

G.193Risk Score

Use only if methodology has:

G.194No Fake Precision

Avoid:

7.43 risk score

without meaningful basis.

G.195Low / Moderate / High

Often enough.

G.196Incident

An incident is:

a defined event that meets established reporting criteria.

G.197Incident Count

Depends on:

G.198More Reports

Can mean:

G.199Interpret Carefully

Always.

G.200Safety Incident

Must be defined by:

G.201Privacy Incident

A privacy incident should mean:

an event involving unauthorized or inappropriate collection, use, access, disclosure, loss, alteration, retention or disposal of personal information, according to the City's applicable privacy framework.

G.202Privacy Concern

May not meet:

Track separately if useful.

G.203Privacy Breach

Use according to:

definition.

G.204Do Not Use "Breach" Casually

No.

G.205Serious Privacy Incident

Should have:

G.206Privacy Incident Count Alone

Not enough.

Also consider:

G.207Cybersecurity Incident

Separate where appropriate.

G.208Cyber Incident Can Become Privacy Incident

Yes.

G.209Accessibility Barrier

An accessibility barrier is:

a physical, digital, communication, procedural, transportation or other obstacle that prevents or materially impairs equitable access to a City service, asset, program or information.

G.210Known Barrier

Identified and recorded.

G.211Removed Barrier

Means:

the barrier itself has been materially corrected or eliminated.

G.212Temporary Accommodation

Not barrier removal.

G.213Mitigation

Not necessarily removal.

G.214Barrier Age

Time since:

G.215Barrier Priority

Based on:

G.216Accessibility Complaint

Separate from:

G.217One Complaint Can Reveal One Systemic Barrier

Yes.

G.218No Complaint Does Not Mean Accessible

No.

G.219Accessibility Review Completed

Activity.

G.220Barrier Removed

Outcome.

G.221Satisfaction

Satisfaction is:

reported resident, user, employee or partner experience measured using a defined survey or feedback method.

G.222Satisfaction Is Subjective

That does not make it:

G.223Satisfaction Is Not Objective Service Performance

Keep separate.

G.224Satisfaction Survey Must State

Population

Method

Dates

Question wording

Sample size

Response rate where known

Weighting where used

Margin of error where statistically appropriate

Major limitations

G.225Open Online Poll

Should not be presented as:

G.226Resident Pulse

Can measure:

G.227Representative Survey

Requires more rigorous:

G.228Self-Selected Survey

Label.

G.229Customer Survey

Measures service users.

Not:

G.230Non-User Views

May differ.

G.231Satisfaction Question Wording

Keep stable.

G.232Changing "Satisfied" to "Very satisfied or satisfied"

Can change results.

Document.

G.233Scale

Define.

Example:

1 to 5.

G.234Net Satisfaction

If used:

Define formula.

G.235Positive Responses

Define.

G.236Neutral Responses

Do not quietly discard without explanation.

G.237No Forced Optimism

Do not word questions to:

G.238No Forced Negativity

Same.

G.239Open Comments

Can add context.

G.240Sentiment Analysis

If AI used:

Treat as:

G.241No Emotion Profiling

Do not classify individual residents by:

G.242Anonymous Feedback

Use where appropriate.

G.243Verified Participation

May be needed for specific civic tools.

G.244Verify Person, Protect Opinion

Apply.

G.245Finance Metric

Financial measures should reconcile to:

G.246Budget

Defined in Appendix E.

G.247Actual

Defined.

G.248Forecast

Defined.

G.249Verified Saving

A verified saving is:

a reduction in actual or credibly unavoidable future municipal cost that has been validated through the City's financial process and is not simply a transfer, deferral, revenue increase or reduction in service disguised as efficiency.

G.250Recurring Verified Saving

Expected to continue in future periods without:

G.251One-Time Saving

Occurs:

G.252Avoided Cost

A future cost credibly prevented.

G.253Capacity Gain

Staff or system capacity released without immediate:

G.254Revenue Increase

Not saving.

G.255Fee Increase

Not saving.

G.256Reserve Draw

Not saving.

G.257Deferred Maintenance

Not saving.

G.258Reduced Service

Not efficiency unless:

support that conclusion.

G.259Tax Requirement

The amount of the municipal budget required from property taxation after other budgeted revenues and financing are taken into account according to the City's adopted financial methodology.

G.260Tax Requirement Is Not

G.261Tax Rate

The rate applied within the applicable property-tax framework.

G.262Individual Tax Bill

Affected by:

G.263Tax Pressure

A plain-language description of:

factors creating upward or downward pressure on the municipal tax requirement or individual bill.

G.264Tax Pressure Is Not One Number Unless Defined

No.

G.265Inflation-Adjusted

A monetary amount converted using a:

G.266Nominal

Actual dollars of the period.

G.267Do Not Mix Nominal and Real

Without disclosure.

G.268Completed Commitment

A commitment is:

Completed

only when the substance of the original public commitment has been delivered according to the criteria established for that commitment.

G.269Announcement Is Not Completion

G.270Funding Application Is Not Completion

G.271Council Approval Is Not Completion

G.272Contract Award Is Not Completion

G.273Construction Start Is Not Completion

G.274Pilot Launch Is Not Permanent Completion

Unless original commitment was:

G.275Completion Criteria

Should be defined when commitment is made or as early as practical.

G.276Completion Can Be Binary

Some commitments are:

G.277Completion Can Be Staged

Others may contain:

G.278Partial Completion

Use:

Partial

where substance is not fully delivered.

G.279Do Not Split Promise Afterward

No.

G.280Example

Promise:

Open a Start-Up Desk.

If desk opens and meets defined basic service:

Could be:

G.281Another Example

Promise:

Build an east-west crossing.

Completing a feasibility study is not:

G.282Missed Commitment

A commitment is:

Missed

when the promised result or milestone was due, has not been delivered, and has not been legitimately reclassified through an openly documented decision before or at the relevant deadline.

G.283Missed Does Not Mean Fraud

Could result from:

G.284Explain Reason

Always.

G.285Stopped

A commitment or initiative may be:

Stopped

when government deliberately decides not to continue.

G.286Responsible Stop

Can be:

G.287But Original Promise Still Exists

Scorecard should show:

G.288Changed

Use when Council openly changes:

before completion.

G.289Changed Is Not Completed

No.

G.290Deferred

Means:

not proceeding on original schedule, but not abandoned.

G.291Deferred Needs

G.292At Risk

Means:

there is a material likelihood that the initiative or commitment will miss its approved scope, schedule, budget, service standard or outcome unless corrective action occurs.

G.293At Risk Is Not Failed

No.

G.294At Risk Should Be Early Warning

Not post-failure label.

G.295On Track

Means:

current evidence reasonably supports delivery within the approved or publicly stated parameters.

G.296On Track Is Forecast Status

Not completed result.

G.297On Hold

Means:

work has deliberately paused pending a defined condition, decision or dependency.

G.298Hold Needs Reason

And:

G.299Cancelled

Use carefully.

Often:

is clearer.

G.300Rejected

Use where decision-making body formally:

G.301Not Started

Different from:

G.302Preparing

Pre-implementation work under way.

G.303Pilot

Temporary test under defined:

G.304Operational

Service or asset is:

G.305Operational Is Not Completed Automatically

A permanent initiative may remain:

rather than "completed."

G.306Completed Project

Can create:

G.307Commitment Status Vocabulary

Use consistently:

Not Started

Preparing

Pilot

Delivering

Operational

Completed

Partial

At Risk

On Hold

Deferred

Stopped

Missed

G.308Do Not Invent Euphemisms

Avoid:

when actual status is:

G.309Traffic-Light System

Traffic lights summarize:

They do not replace:

G.310Green

Generally means:

meeting the defined standard, target or expected range based on sufficiently reliable information.

G.311Amber

Generally means:

material concern, deterioration, uncertainty or risk requiring attention but not yet meeting the Red definition.

G.312Red

Generally means:

defined standard materially missed, serious risk present, or corrective decision required.

G.313Grey

Means:

insufficient verified information to assign Green, Amber or Red responsibly.

G.314Grey Is Not Neutral

It means:

G.315Grey Is Not Zero

Never.

G.316Grey Can Be Serious

Especially where:

are involved.

G.317Metric-Specific Thresholds

Green, Amber and Red should not use:

G.318Define Thresholds

For each headline measure where traffic lights are used.

G.319No Changing Threshold After Result

Never.

G.320Threshold Revision

If needed:

G.321Red Does Not Mean Blame

A Red result may reflect:

G.322Explain Owner

Separate:

G.323Green Does Not Mean Permanent Success

Conditions can change.

G.324One Green Quarter

Not proof of:

G.325Trend

A trend is:

direction across a sufficient series of comparable observations.

G.326Two Points

May show change.

Not always:

G.327Trend Period

Define.

G.328Seasonal Trend

Use comparable periods.

G.329Rolling Average

Can smooth:

G.330Rolling Period

Define.

G.331Year-to-Date

Define start.

G.332Fiscal Year

Use official City fiscal period.

G.333Calendar Year

Different.

G.334Election Term

Different.

G.335Do Not Mix Periods

Without explanation.

G.336Data Owner

The institutional role accountable for:

the source data.

G.337Metric Owner

The role responsible for:

G.338Data Owner and Metric Owner May Differ

Yes.

G.339Independent Verification

Use where:

measure warrants it.

G.340Finance Verification

For:

G.341Engineering Verification

For:

G.342Accessibility Verification

May require:

review.

G.343Privacy Verification

May require:

review.

G.344Survey Verification

May require:

G.345Self-Reported Department Metric

Can be valid.

Label source.

G.346Evidence Quality

Possible categories:

Verified Administrative Data

Professional Assessment

Survey

Sample

Estimate

Model

Self-Reported Partner Data

External Data

G.347Confidence

Can summarize source quality.

G.348Partner-Reported Data

Do not automatically present as:

G.349Partner Reporting

State:

Reported by partner

where appropriate.

G.350External Government Data

State source.

G.351Data Cut-Off

Each report should identify:

Data current to [date].

G.352Publication Lag

Explain if significant.

G.353Revision

Data may change after:

G.354Preliminary Results

Label.

G.355Restatement

A restatement is:

a revised historical result caused by corrected data, methodology change or improved information.

G.356Restatement Does Not Erase Original

Maintain:

G.357Material Restatement

Explain:

What changed?

Why?

Which periods?

Effect on trend?

G.358Minor Correction

Can use proportionate treatment.

G.359Methodology Change

Should state:

Old method

New method

Reason

Effective date

Historical restatement available?

G.360Comparable Series

If historical data can be recalculated:

Prefer.

G.361Break in Series

If not:

Show.

G.362Do Not Draw Trend Across Non-Comparable Break

No.

G.363Data Missing

Use:

Not Available

or:

Unknown

according to context.

G.364Zero

Means:

verified count or value of zero.

G.365Zero Is Data

Unknown is not.

G.366No Response

Different from:

G.367Suppressed

Use where data exists but cannot be responsibly published due to:

G.368Suppressed Is Not Unknown

Internally known.

Publicly withheld.

G.369Not Applicable

Means measure legitimately does not apply.

G.370Do Not Use N/A to Hide Missing Data

No.

G.371Data Quality

A headline metric should disclose known limitations.

G.372Data Quality Dimensions

Possible:

G.373Data Quality Score

Not required.

G.374Plain-Language Limitation

Often better.

G.375Sample Size

Important for:

G.376Small Sample Warning

Use.

G.377Sampling Method

Explain.

G.378Convenience Sample

Label.

G.379Random Sample

Label if genuinely random.

G.380Census of Records

Use when all applicable records included.

G.381Weighting

If survey weighted:

Explain.

G.382Margin of Error

Use only where statistically valid.

G.383Do Not Put Margin of Error on Self-Selected Poll

No.

G.384Privacy in Measurement

Collect:

G.385Scorecard Is Not Reason to Build Resident Profiles

No.

G.386Aggregation

Prefer where individual identity unnecessary.

G.387Small Cell Suppression

Use where re-identification risk exists.

G.388Location Granularity

Do not publish vulnerable-person heat maps.

G.389Youth

Use stronger privacy standards.

G.390Health

Use stronger privacy standards.

G.391Political Opinion

Do not attach to identifiable resident performance record.

G.392Faith

Same.

G.393Disability

Do not collect diagnosis merely to measure accessibility unless genuinely necessary and lawful.

G.394Participation

Can often be measured without:

G.395Unique Participation

If exact identity is unnecessary:

Use privacy-preserving methodology.

G.396Data Linkage

Combining datasets can create:

No.

G.398AI and Scorecards

AI can assist with:

G.399AI Cannot Create Facts

No.

G.400AI-Generated Measure

Must still have:

G.401AI Inference

Label if result is:

G.402Sentiment

Do not use AI sentiment as:

G.403Emotion Detection

Do not use.

G.404Automated Risk Scoring of Residents

Do not use for public Scorecard.

G.405Employee Ranking

Scorecards should generally measure:

Not create simplistic public rankings of:

G.406Department Comparison

Can be useful.

But services differ.

G.407No Leaderboard Theatre

No.

G.408Finance Scorecard Definitions

Core measures may include:

Operating Budget

Actual Operating Result

Forecast

Debt

Debt Service

Reserves

Verified Savings

Capital Commitments

Unfunded Mandates

All should follow Appendix E.

G.409Tax Scorecard Definitions

Core distinctions:

Tax Requirement

Tax Rate

Individual Bill

County Component

Education Component

Fees

Assessment

Do not merge.

G.410Service Scorecard Definitions

Core concepts:

Request

Acknowledgement

Meaningful Response

First-Contact Resolution

Transfer

Handoff

Total Elapsed Time

City-Controlled Time

Backlog

Reopen

Repeat Issue

Closure

G.411Infrastructure Scorecard Definitions

Core concepts:

Asset

Condition

Confidence

Criticality

Inspection

Preventive Maintenance

Deferred Maintenance

Capital Need

Completion

Service Failure

G.412Downtown Scorecard Definitions

Core concepts may include:

Storefront

Vacant Storefront

Occupied Storefront

Upper-Floor Unit / Space

Opening

Closure

Pedestrian Passage

Parking Occupancy

Public-Realm Issue

Ordinary-Day Measurement

G.413Storefront Vacancy

Must have:

G.414No "Looks Empty"

Use defined method.

G.415Business Opening

Define whether:

G.416Business Closure

Likewise.

G.417Relocation Within City

Should not automatically count as:

without context.

G.418Pedestrian Passage

A counter event is:

a passage past a counter.

Not necessarily:

G.419Housing Scorecard Definitions

Core stages:

Inquiry

Pre-Application

Formal Application

Approval

Permit

Start

Completion

Occupancy

G.420Approval Is Not Home

Always.

G.421Permit Is Not Start

G.422Start Is Not Completion

G.423Completion Is Not Necessarily Occupancy

G.424Net New Unit

Should account for:

according to defined methodology.

G.425Affordable

Must always specify:

G.426Accessible Unit

Must specify actual:

G.427Business Scorecard Definitions

Core concepts:

Inquiry

Start-Up Desk User

Business Opening

Closure

Expansion

Relocation

Licence

Permit

Procurement Bidder

Local Vendor

Canadian Supplier

Verified Public Incentive

G.428Local Business

Define geographic boundary.

G.429Local Supplier

Same.

G.430Ontario Supplier

Different.

G.431Canadian Supplier

Different.

G.432Do Not Change Definition Depending on Result

No.

G.433Safety Scorecard Definitions

Core concepts:

Harm

Incident

Call for Service

Response

Serious Injury

Repeat Demand

Capacity

Prevention Activity

Perception of Safety

G.434Call for Service

Activity / demand measure.

Not automatic:

G.435Arrest

Enforcement activity.

Not automatic:

G.436Ticket

Same.

G.437Perceived Safety

Survey measure.

G.438Actual Harm

Administrative or professional measure.

Keep separate.

G.439Right Responder

Success should examine:

Not just:

G.440Recreation Scorecard Definitions

Core concepts:

Registration

Attendance

Unique Participant

Program

Session

Cancellation

Waitlist

Facility Availability

Utilization

Equipment Loan

Beginner Opportunity

Adaptive Opportunity

G.441Program Count

Not outcome.

G.442Facility Hours Available

Different from:

G.443Facility Utilization

Define available denominator.

G.444Planned Closure

Separate from:

G.445Equipment Loan

Transaction count.

Not unique users.

G.446Youth Scorecard Definitions

Core concepts:

Opportunity

Applicant

Participant

Paid Role

Volunteer Role

Co-op Role

Training

Certification

Mentor

Completed Project

Next Step

G.447Opportunity

A genuinely available role or activity with:

G.448Listing Is Not Filled Opportunity

No.

G.449Paid Role

Must be genuinely:

G.450Honorarium

Do not automatically classify as:

G.451Volunteer

Must be genuinely:

G.452Certification

Define issuing:

G.453Next Step

Could include:

Do not claim City caused it without evidence.

G.454Accessibility Scorecard Definitions

Core concepts:

Barrier

Barrier Removed

Temporary Accommodation

Accessible Journey

Digital Accessibility Issue

Winter Accessibility Issue

Accessibility Review

Repeat Barrier

G.455Accessibility Review

Activity.

G.456Barrier Removal

Outcome.

G.457Accessible Journey

Includes all relevant stages from:

G.458Privacy and Information Scorecard Definitions

Core concepts:

Personal Information

Data Collection

Data Sharing

Retention

Access

Privacy Incident

Correction

Official Information

Tracker

Digital System

Exit Test

G.459Official Information

Information published by the City through an authorized:

G.460Correction

A material change made because previous official information was:

G.461Update

A current fact changed.

Not same as:

G.462Data Collection

Information newly obtained.

G.463Data Access

Someone views or retrieves information.

Different.

G.464Data Sharing

Information provided to another:

outside the original operational boundary according to defined rules.

G.465Environment Scorecard Definitions

Core concepts may include:

Water Loss

Wastewater Incident

Flooding Issue

Tree Planted

Tree Survived

Energy Use

Waste Generated

Diversion

Asset Resilience

Environmental Liability

G.466Tree Survived

Define observation period.

G.467One-Month Survival

Not meaningful long-term tree success.

G.468Waste Diversion

Define:

G.469Energy Reduction

Needs:

G.470GHG Estimate

If used:

State methodology and uncertainty.

G.471Partnership Scorecard Definitions

Core concepts:

Partnership

Government-to-Government Relationship

Shared Service

Grant

Contract

Sponsorship

Referral

Commitment

Deliverable

Handoff

External Contribution

Exit

G.472Meeting

Activity.

G.473MOU

Output.

G.474Better Resident Result

Outcome.

G.475Partner Count

Not success by itself.

G.476External Contribution

Separate:

G.477Leveraged Money

Use only with:

G.478SON Relationship Measures

Should measure:

Do not score:

G.479Ontario Relationship Measures

Track:

Not just:

G.480Federal Relationship Measures

Same.

G.481County Relationship Measures

Track:

G.482Satisfaction Scorecard Definitions

Should distinguish:

Overall Resident Satisfaction

Service-Specific Satisfaction

User Satisfaction

Trust

Confidence

Perceived Value

Perceived Safety

Ease of Access

G.483Trust and Satisfaction Are Not Same

No.

G.484Satisfaction With Decision

Different from:

G.485Resident Can Dislike Decision and Still Rate Process Fairly

Important.

G.486Resident Can Like Outcome and Still Experience Poor Process

Also.

G.487Completed Scorecard Definitions

Each commitment should have:

Original wording

Date made

Owner

Due date

Completion criteria

Evidence

Completion date

Status

G.488Missed Scorecard Definitions

Each missed commitment should show:

Original wording

Due date

What was delivered

What was not delivered

Reason

Cost to date

Next status

G.489No Promise Deletion

Ever.

G.490Commitment Version

If commitment legitimately changes:

Preserve:

G.491Election Promise Versus Council Commitment

Different.

G.492Campaign Promise

Political commitment by candidate or campaign.

G.493Council Commitment

Institutional commitment adopted through:

where applicable.

G.494Mayor Commitment

Could remain personal political commitment even if Council:

G.495Score Accordingly

Do not mark Council:

for rejecting Mayor proposal.

G.496Mayor Promise Requiring Council

Status should state:

Council approval required.

G.497Promise Requiring Ontario

State:

Ontario action required.

G.498Promise Requiring Canada

State.

G.499Promise Requiring County

State.

G.500Accountability Requires Jurisdiction

Always.

G.501Measurement Frequency

How often data are:

G.502Reporting Frequency

How often data are:

G.503They Can Differ

Yes.

G.504Real-Time Dashboard

Not always useful.

G.505Monthly

Good for some services.

G.506Quarterly

Good for many public measures.

G.507Annual

Good for slower outcomes.

G.508Event-Based

Good for:

G.509Do Not Over-Report Noise

No.

G.510Do Not Under-Report Problems

No.

G.511Reporting Lag

State.

G.512Data Cut-Off Consistency

Important.

G.513Quarterly Comparison

Use same:

G.514Leap Year

Usually immaterial.

But exact daily rates may need:

G.515Partial Period

Label.

G.516Annualization

A partial-period value projected to a full year.

G.517Annualized Is Forecast

Not actual.

G.518Year-over-Year

Compare same period.

G.519Quarter-over-Quarter

Can be distorted by:

G.520Trend Commentary

Should distinguish:

G.521Correlation

Not causation.

G.522Public Narrative

Every Scorecard should pair numbers with short:

G.523Explanation Should Include

What changed?

Why might it have changed?

Is cause known?

What is being done?

G.524Do Not Spin Red Green

No.

G.525Do Not Spin Green Red

Also.

G.526Context Is Not Excuse

It is:

G.527Benchmark

A comparator such as:

G.528Peer Comparison

Use carefully.

G.529Peer Municipality

Should be sufficiently comparable for:

G.530Population Alone Does Not Make Peer

No.

G.531Service Model

Consider.

G.532Two-Tier Structure

Consider.

G.533Geography

Consider.

G.534Asset Age

Consider.

G.535Tourism

Consider.

G.536External Benchmark Source

Identify.

G.537Benchmark Year

Identify.

G.538Benchmark Does Not Replace Local Target

No.

G.539Rank

Use cautiously.

G.540"Number One"

Depends on:

G.541No Cherry-Picked Ranking

Never.

G.542Cost Per Unit

Can be useful.

G.543Unit Must Be Comparable

G.544Lower Cost Per Unit

Not necessarily better if:

G.545Pair Cost and Outcome

Where possible.

G.546Productivity

Output per unit of:

according to defined method.

G.547Productivity Is Not Work Intensity

Do not reward:

G.548Service Quality

Needs separate measure.

G.549Quality Error

Could include:

G.550Quality Definition

Service-specific.

G.551Reliability

The extent service performs:

G.552Uptime

Useful for:

G.553Planned Downtime

Separate.

G.554Unplanned Downtime

Separate.

G.555Uptime Percentage

Needs defined:

G.55624/7 Service

Different denominator from:

G.557Cancellation

For recreation or public programs:

Define.

G.558Weather Cancellation

Separate where useful.

G.559Staffing Cancellation

Different diagnostic cause.

G.560Waitlist

A count of:

according to defined criteria.

G.561Duplicate Waitlist Entries

Remove appropriately.

G.562Waitlist Demand

Not same as:

Some people do not join waitlists.

G.563Vacancy

For downtown:

Define physically.

For staffing:

Different definition.

G.564Same Word, Different Metric

Metric card prevents confusion.

G.565Completion Rate

Completed Eligible Items ÷ Total Eligible Items

But define:

G.566Carryover

Important.

G.567Cohort Method

Sometimes needed.

G.568Example

Youth program completion should follow:

not all people who ever expressed interest.

G.569Conversion Rate

Movement from one stage to next.

G.570Housing Conversion

Examples:

G.571Conversion Rate Requires Cohort Discipline

Do not mix unrelated periods blindly.

G.572Pipeline

A sequence of defined stages.

G.573Pipeline Count

Snapshot of items currently at each:

G.574Flow

Items moving through stage during:

G.575Snapshot and Flow Are Different

Do not mix.

G.576Point-in-Time

Value at specific:

G.577Period Measure

Value over:

G.578Cumulative

Total since defined:

G.579Cumulative Measures Can Hide Slowdown

Pair with current-period data.

G.580Four-Year Total

Useful.

G.581Annual Breakdown

Also.

G.582Denominator Freeze

For some four-year commitments:

Use original denominator if the question is:

How much of the original backlog was cleared?

G.583Dynamic Denominator

For other measures:

New items should enter denominator.

G.584Choose Before Reporting

Critical.

G.585Backlog Reduction Example

If backlog starts at 1,000 and 800 are completed but 600 new overdue items appear:

Do not claim:

80% backlog eliminated

without showing:

G.586Both Views

Show:

Original backlog resolved

Current backlog

G.587Cohort Survival

Useful for:

where appropriate.

G.588Tree Survival Cohort

Trees planted in same season followed over:

G.589Business Survival

External economic outcome.

Do not claim City caused it.

G.590Program Retention

Could be useful.

G.591Data Integrity

The City should maintain controls against:

G.592Audit Trail

For important Scorecard data.

G.593Manual Override

Should be:

G.594Data Correction

Document.

G.595Spreadsheet Dependency

Can be practical.

But critical metrics should avoid:

dependency.

G.596Reproducibility

A reasonably skilled successor should be able to:

reproduce the reported number from the source and method.

G.597Metric Documentation

Essential.

G.598Bus Factor

Do not allow one employee to be the only person who knows:

G.599Code

If calculations use code:

G.600Formula

Publish where appropriate.

G.601Proprietary Vendor Metric

Do not rely on score you:

G.602Black-Box Score

Avoid for public accountability.

G.603Vendor Dashboard

Can supplement.

Not become official truth automatically.

G.604Raw Data

May not be publicly releasable.

G.605Methodology Can Still Be Public

Often.

G.606Auditability

Public Scorecards should be:

G.607Source Record

Preserve according to:

G.608Public Archive

Past Scorecards should remain:

G.609Do Not Replace Old Dashboard

Archive snapshots.

G.610Historical Accountability

Residents should be able to see:

G.611Election-Year Rule

Do not change headline metric methodology shortly before election unless:

G.612Necessary Change

Publish:

G.613Campaign Firewall

Official Scorecard remains:

G.614Campaign May Cite It

Like any member of public.

G.615Campaign Cannot Edit Official Scorecard

No.

G.616Mayor Cannot Suppress Red Result

No.

G.617Council Cannot Vote to Recolour Metric

No.

G.618Professional Rules Apply

Thresholds established:

G.619Scorecard Governance

The City should identify:

Metric owner

Data owner

Publication owner

Verification owner

G.620Finance Measures

Finance verifies.

G.621Technical Measures

Qualified technical owner verifies.

G.622Corporate Reporting

Can coordinate.

G.623Council

Receives and questions.

G.624Mayor

Explains and leads response.

G.625Mayor Does Not Rewrite Data

No.

G.626Public Challenge

Residents should be able to:

G.627Correction Route

Provide.

G.628Valid Challenge

Review.

G.629Frivolous Challenge

Does not require changing measure.

G.630Methodology Notes

Publicly accessible.

G.631Plain-Language First

Technical appendix available where needed.

G.632Scorecard Page Structure

Each public Scorecard page should ideally include:

Question

Current Result

Baseline

Previous Period

Target / Standard

Status

Trend

Definition

Source

Last Updated

Explanation

Next Action

G.633What Is Missing?

Also useful.

G.634Confidence

Where appropriate.

G.635Grey Explanation

If unknown:

Explain:

G.636Data Improvement Is an Outcome

Improving from:

to:

can itself be valuable.

G.637But Do Not Celebrate Measurement as Service Outcome

No.

G.638Measuring Potholes

Not same as:

G.639Measuring Accessibility Barriers

Not same as:

G.640Measuring Privacy Risk

Not same as:

G.641Measurement Burden

Data collection costs:

G.642Proportionality

Do not create expensive data system for:

G.643Use Existing Administrative Data

Where reliable.

G.644Do Not Collect Personal Data Just Because Metric Would Be Interesting

No.

G.645Manual Sampling

Can sometimes be:

G.646Temporary Measurement

Can support:

G.647Retire Metric

A metric can be retired when:

G.648Metric Retirement

Record:

G.649Do Not Retire Because Result Is Bad

Never.

G.650Replacement Metric

Explain relationship.

G.651Scorecard Review

Annually ask:

Does this metric still answer an important question?

Is definition still valid?

Is data reliable?

Is reporting burden justified?

Is it being gamed?

Is a better measure available?

G.652Metric Portfolio

Keep lean.

G.653Outcome Priority

Headline measures should lean toward:

with enough activity measures to explain them.

G.654Control

Measures should distinguish:

City Controlled

Shared

External Context

G.655City Controlled

City can directly influence core result.

G.656Shared

Result depends materially on:

G.657External Context

City largely cannot control.

G.658Example

Housing approvals:

Housing starts:

Interest rates:

G.659Scorecard Should Not Punish City for Weather

But City should be measured on:

G.660Economic Conditions

Context.

G.661Population Change

Context.

G.662Provincial Law Change

Context.

G.663Federal Program Change

Context.

G.664Context Does Not Erase Responsibility

City-controlled part remains.

G.665Responsible Owner

For shared measures:

Identify City's:

G.666External Dependency

Identify.

G.667Partner Commitment

Identify where formal.

G.668Score City Commitments

Not another government as if subordinate.

G.669SON

Especially.

G.670No Nation Score

Never produce:

SON performance score.

G.671Score Relationship Process

Where mutually appropriate.

G.672Score City's Own Commitments

Yes.

G.673County

Can compare shared commitments.

G.674Ontario

Track response without pretending City controls it.

G.675Canada

Same.

G.676Public Narrative Standard

Use verbs carefully:

Delivered

Approved

Funded

Started

Completed

Reduced

Increased

Contributed

Coincided

G.677"Delivered"

Should identify:

G.678"Created"

Use only where City actually:

G.679"Saved"

Only verified under Appendix E.

G.680"Reduced"

Requires baseline.

G.681"Improved"

Requires defined comparison.

G.682"Record"

Requires comparable historical data.

G.683"Best Ever"

Very high evidentiary threshold.

G.684"First"

Verify.

G.685"Number One"

Verify comparator.

G.686"Free"

Avoid when taxpayer or partner pays.

G.687"Fully Funded"

Define funding sources and operating implications.

G.688"Affordable"

Define.

G.689"Accessible"

Define.

G.690"Safe"

Avoid absolute claim.

G.691"Secure"

Avoid absolute claim.

G.692"Private"

Explain actual privacy design.

G.693"Anonymous"

Use only if re-identification risk is appropriately addressed.

G.694"Open"

Explain access.

G.695"Canadian"

Define if relevant to supplier or hosting.

G.696"Local"

Define boundary.

G.697"Completed"

Use the completion standard.

G.698Anti-Gaming Rule One

Do not change a definition because:

G.699Rule Two

Do not change a denominator after seeing the result.

G.700Rule Three

Do not move baseline forward to erase weak early performance.

G.701Rule Four

Do not choose favourable seasonal period after results appear.

G.702Rule Five

Do not call Unknown:

G.703Rule Six

Do not call suppressed data:

G.704Rule Seven

Do not call preliminary:

G.705Rule Eight

Do not call estimate:

G.706Rule Nine

Do not call forecast:

G.707Rule Ten

Do not call target:

G.708Rule Eleven

Do not call activity:

G.709Rule Twelve

Do not call output:

G.710Rule Thirteen

Do not call outcome:

without evidence.

G.711Rule Fourteen

Do not call correlation:

G.712Rule Fifteen

Do not count registrations as:

G.713Rule Sixteen

Do not count visits as:

G.714Rule Seventeen

Do not create unnecessary identity tracking merely to obtain unique counts.

G.715Rule Eighteen

Do not count automated acknowledgement as:

G.716Rule Nineteen

Do not close unresolved files solely to improve:

G.717Rule Twenty

Do not clear backlog by:

without legitimate reason.

G.718Rule Twenty-One

Do not pause service clock while file remains:

without defined pause rule.

G.719Rule Twenty-Two

Do not exclude difficult cases after:

G.720Rule Twenty-Three

Do not include easy cases selectively to improve:

G.721Rule Twenty-Four

Do not report only mean when long-tail failure matters.

G.722Rule Twenty-Five

Do not choose percentile after seeing:

G.723Rule Twenty-Six

Do not count complaint reduction as success if complaint access became:

G.724Rule Twenty-Seven

Do not call zero reported incidents:

unless detection and reporting are reliable.

G.725Rule Twenty-Eight

Do not call temporary accessibility accommodation:

G.726Rule Twenty-Nine

Do not call accessibility review:

G.727Rule Thirty

Do not call tree planting:

G.728Rule Thirty-One

Do not call grant announcement:

G.729Rule Thirty-Two

Do not call forecast saving:

G.730Rule Thirty-Three

Do not call capacity gain:

G.731Rule Thirty-Four

Do not call tax-rate freeze:

G.732Rule Thirty-Five

Do not call City tax change:

without relevant context.

G.733Rule Thirty-Six

Do not call an approved housing unit:

G.734Rule Thirty-Seven

Do not call building permit:

G.735Rule Thirty-Eight

Do not call housing start:

G.736Rule Thirty-Nine

Do not count a business relocation within City as both a net new business and closure without explaining.

G.737Rule Forty

Do not call enforcement activity:

G.738Rule Forty-One

Do not call meeting count:

G.739Rule Forty-Two

Do not call MOU:

G.740Rule Forty-Three

Do not score another government or Indigenous Nation as if it reports to Owen Sound.

G.741Rule Forty-Four

Do not call survey participants:

G.742Rule Forty-Five

Do not attach statistical margin of error to methodology that does not support it.

G.743Rule Forty-Six

Do not change survey wording and compare results as if identical.

G.744Rule Forty-Seven

Do not use AI sentiment analysis as representative resident opinion.

G.745Rule Forty-Eight

Do not use resident sentiment or behaviour scoring.

G.746Rule Forty-Nine

Do not rank individual employees publicly using simplistic Scorecard numbers.

G.747Rule Fifty

Do not change Green, Amber or Red thresholds after:

G.748Rule Fifty-One

Do not colour Unknown:

G.749Rule Fifty-Two

Do not recolour Red through Council vote.

G.750Rule Fifty-Three

Do not remove poor metrics during election year.

G.751Rule Fifty-Four

Do not introduce favourable metrics during election year without preserving comparability.

G.752Rule Fifty-Five

Do not remove original public commitments after they are:

G.753Rule Fifty-Six

Do not call study completion:

unless study itself was the commitment.

G.754Rule Fifty-Seven

Do not split one promise into multiple completed commitments after the fact.

G.755Rule Fifty-Eight

Do not combine several missed promises into one vague miss.

G.756Rule Fifty-Nine

Do not mark commitment completed merely because:

G.757Rule Sixty

Do not mark commitment missed merely because:

if original result was actually delivered.

G.758Original Substance Matters

Not wording manipulation.

G.759Measurement Integrity Test

For every headline measure ask:

Would the definition be the same if the result were bad?

G.760Baseline Test

Was the baseline selected before we knew which comparison looked best?

G.761Denominator Test

Did the denominator remain consistent?

G.762Source Test

Can we identify where the number came from?

G.763Reproducibility Test

Could someone else reasonably reproduce it?

G.764Confidence Test

How sure are we?

G.765Missing-Data Test

Are unknowns shown as unknown?

G.766Privacy Test

Could we measure this with less personal information?

G.767Accessibility Test

Can residents understand the measure?

G.768Causation Test

Are we claiming more credit or blame than the evidence supports?

G.769Target Test

Was the target established before the result?

G.770Traffic-Light Test

Were thresholds established before the colour was known?

G.771Completion Test

Did we deliver the original substance of the promise?

G.772Miss Test

Was the commitment due, and was it actually not delivered?

G.773Stop Test

Did we preserve the original record even after deciding to stop?

G.774Election Test

Would this metric still be reported the same way if another Mayor were running for re-election?

G.775Reverse Test

Would we accept this measurement method if it made a political opponent look good?

G.776Public Understanding Test

Can an ordinary resident understand what the number does and does not mean?

G.777Scorecard Publication Standard

At minimum, every headline metric should publish:

Current result

Baseline

Definition

Source

Period

Status

Trend where useful

Target or standard where applicable

Limitations

Last updated

Provide where detail needed.

G.779No Raw Data Requirement

Public should not need:

to understand headline result.

G.780Open Data

Where lawful and useful:

Can support independent analysis.

G.781Privacy and Security

Still apply.

G.782Machine-Readable

Useful where appropriate.

G.783Human-Readable

Essential.

G.784Print

Provide core Scorecard in a form that can be:

G.785No App Required

Ever for public access to essential Scorecard information.

G.786Accessibility

Public Scorecard should be:

G.787Colour Is Supplemental

Do not rely on:

G.788Label

Use:

or comparable text.

G.789Charts

Should include:

G.790Truncated Axis

Use carefully.

G.791Visual Manipulation

Avoid.

G.792Scale Changes

Disclose.

G.793Map Measures

Maps should not reveal:

G.794Geographic Aggregation

Use safe areas.

G.795Neighbourhood Comparison

Can be useful for:

G.796Do Not Rank Neighbourhoods as "Good" and "Bad"

No.

G.797Equity Analysis

Can examine whether service barriers are distributed:

G.798Identity Data

Use only where lawful and necessary.

G.799Geographic Proxy

Do not use carelessly to infer:

G.800Scorecard Governance Committee

The City may assign corporate oversight to an existing:

structure.

Do not create a new committee merely because:

G.801Independent Review

Consider periodic review by:

for selected high-risk metrics.

G.802Not Every Metric Needs External Audit

Proportionality.

G.803Financial Measures

Higher verification standard.

G.804Election Commitments

Higher transparency standard.

G.805Privacy Measures

Higher sensitivity.

G.806Critical Infrastructure

Higher technical standard.

G.807Public Survey

Higher methodological standard where representative claims are made.

G.808Year One

Establish:

G.809Year One Priority

Do not rush to produce a polished dashboard with:

G.810Definitions Before Graphics

Always.

G.811Year Two

Improve:

G.812Year Three

Review:

G.813Retire Weak Metrics

Openly.

G.814Add Better Metrics

Openly.

G.815Preserve Series Breaks

G.816Year Four

Publish:

Four-Year Measurement Integrity Audit.

G.817Four-Year Audit

Should answer:

Which measures remained unchanged for four years?

Which definitions changed?

Why?

Which baselines were revised?

Which data were restated?

Which metrics were retired?

Which were added?

Where did data quality improve?

Where did unknowns remain?

Which measures proved vulnerable to gaming?

Which outcome measures improved?

Which did not?

Which public claims required correction?

G.818Name the Best Metric

The one that most improved:

G.819Name the Weakest Metric

And explain:

G.820Name the Metric Retired

If any.

G.821Name the Largest Restatement

If any.

G.822Name the Largest Baseline Error Corrected

If any.

G.823Name the Most Important Unknown Converted to Known

G.824Name the Largest Remaining Data Gap

G.825Name the Measure Most Often Misunderstood

G.826Name the Measure Most Vulnerable to Gaming

G.827Name a Metric Whose Result Was Worse Than Expected

G.828Name a Metric Whose Result Was Better Than Expected

G.829Name a Forecast That Was Wrong

Where material.

G.830Name an Activity That Did Not Produce Intended Outcome

G.831Name an Outcome Achieved With Less Activity Than Expected

G.832Name a Metric Changed Because Privacy Review Found Excessive Data Collection

If applicable.

G.833Name a Metric Improved Because Accessibility Made Reporting More Representative

If applicable.

G.834Name a Public Claim Corrected Through Scorecard Review

G.835Election-Year Integrity

Publish final pre-election Scorecard using:

G.836Do Not Freeze Bad Data

Corrections remain allowed.

G.837But Preserve Audit Trail

Always.

G.838Next-Council Handoff

Provide:

Metric dictionary

Formulas

Baselines

Targets

Source systems

Data owners

Methodology history

Known limitations

Open quality issues

Scheduled reporting dates

G.839No Institutional Memory Trap

The Scorecard should survive:

turnover.

G.840Definitions Belong to Institution

Not:

G.841The Scorecard Definitions Commitment

Owen Sound should commit to:

Define every recurring headline metric before using it to judge performance.

Maintain a Metric Definition Card for each major public measure.

Give every metric a plain-language question so residents understand why it exists.

Define numerator, denominator, inclusion, exclusion, data source, owner, period and limitations where applicable.

Establish baselines before interventions wherever reasonably possible.

Use Unknown rather than inventing historical precision when no reliable baseline exists.

Never quietly change a baseline after results are known.

Distinguish targets, forecasts, estimates and actual results.

Do not present targets as guarantees.

Do not create arbitrary targets merely because round numbers appear ambitious.

Use target ranges where they better reflect good performance.

Distinguish activity, output, outcome and longer-term impact.

Do not claim causal impact where the evidence supports only contribution or coincidence.

Define counts precisely.

Separate registrations, attendance, completion, repeat participation and unique participants.

Do not build unnecessary identity systems merely to obtain perfect unique counts.

Define every rate with an explicit numerator and denominator.

Distinguish percentage changes from percentage-point changes.

Define per-capita, per-household, per-kilometre and per-application measures carefully.

Do not mix centreline kilometres and lane-kilometres.

Use mean, median and long-tail percentiles according to what they reveal.

Define percentile measures before results are known.

Use small-sample warnings where appropriate.

Give all time measures explicit clock-start, clock-stop and pause rules.

Separate total elapsed time from City-controlled time where the distinction matters.

Do not count automated receipts as meaningful responses unless the metric is specifically acknowledgement time.

Define meaningful response, closure, reopen and repeat issue consistently.

Use legitimate closure codes.

Never close files merely to improve service statistics.

Define backlog separately from ordinary open work.

Track backlog age and high-priority backlog where useful.

Do not reduce backlog through deletion or reclassification designed to improve the number.

Define service transfers, external handoffs and warm handoffs.

Do not reduce transfer counts by simply rejecting more requests.

Pair handoff metrics with resolution, repeat contact and resident experience.

Define asset condition separately by appropriate asset class.

Publish condition confidence where it materially affects interpretation.

Never treat Unknown condition as Good.

Keep condition, confidence, criticality and risk separate.

Define incidents according to clear reporting criteria.

Recognize that more reported incidents can reflect better detection or reporting rather than more underlying harm.

Define privacy incidents carefully and distinguish them from concerns and cybersecurity incidents.

Measure privacy incident severity and recurrence rather than count alone.

Define accessibility barriers broadly enough to include physical, digital, communication, transportation and procedural barriers.

Do not call temporary accommodation barrier removal.

Do not call accessibility review an accessibility outcome.

Define satisfaction as a subjective experience measure and keep it separate from objective service performance.

Publish survey population, method, dates, wording and sample size.

Never present a self-selected online poll as representative of the whole population.

Keep satisfaction questions stable over time where trend comparison is intended.

Do not use AI sentiment analysis as a substitute for properly designed resident research.

Never use emotion or personality profiling of residents as a civic performance measure.

Define verified savings through the financial framework.

Never classify revenue increases, fee increases, reserve draws, deferred maintenance or capacity gains as cash savings.

Define tax requirement separately from tax rate and individual tax bill.

Keep City, County and education tax components distinguishable.

Define Completed according to the original substance of the commitment.

Do not call announcements, approvals, funding applications, contract awards or project starts completed unless that was the actual commitment.

Use Partial when a commitment has only been partly delivered.

Define Missed clearly and publish the reason.

Allow responsible stops without deleting the original commitment.

Use Changed, Deferred, On Hold, At Risk and Stopped according to defined rules rather than political convenience.

Use one City-wide commitment-status vocabulary.

Do not invent euphemisms for missed or delayed commitments.

Use Green, Amber, Red and Grey only according to pre-established metric-specific thresholds.

Define Grey as insufficient verified information.

Never colour Unknown Green.

Never change traffic-light thresholds after the result is known.

Recognize that Red identifies a condition requiring attention, not necessarily a person to blame.

Define trends using sufficient comparable observations.

Treat seasonal measures with appropriate seasonal comparison.

Distinguish year-to-date, fiscal-year, calendar-year, rolling and election-term results.

Assign every headline metric a data owner and metric owner.

Use independent or professional verification where financial, technical, privacy, legal or statistical risk warrants it.

Label partner-reported and externally reported data rather than presenting everything as City-verified.

Publish data cut-off dates and reporting lags.

Use preliminary labels before reconciliation is complete.

Maintain a transparent restatement policy.

Preserve original reported results when material data or methodology corrections occur.

Publish old method, new method, reason, effective date and historical comparability when methodology changes.

Do not draw continuous trend lines across non-comparable methodology breaks without warning.

Distinguish Zero, Unknown, Suppressed and Not Applicable.

Never use zero as a placeholder for missing data.

Publish known data-quality limitations.

Use representative sampling claims only when the sampling method supports them.

Do not publish statistical margins of error for surveys where the methodology does not justify them.

Collect the least personal information necessary to produce useful Scorecard measures.

Use aggregation and small-cell suppression where needed to protect privacy.

Never build vulnerable-person heat maps for public performance reporting.

Apply stronger privacy protection to youth, health, disability, political and faith-related information.

Do not link datasets merely because a combined metric would be interesting.

Use AI only as an assistive measurement tool subject to human verification.

Do not allow AI-generated summaries to become the source of record.

Do not use AI to create resident risk, trustworthiness or civic-value scores.

Do not turn public Scorecards into simplistic individual employee rankings.

Measure Finance, Tax, Services, Infrastructure, Downtown, Housing, Business, Safety, Recreation, Youth, Accessibility, Privacy, Environment, Partnership, Satisfaction, Completed and Missed commitments according to the specialized definitions in this appendix.

For housing, preserve the pipeline from inquiry through occupancy and never count approval as a home.

For downtown, use a stable geographic boundary and distinguish physical storefront status from impressions.

For business, define local, Ontario and Canadian suppliers separately.

For safety, distinguish harm, demand, enforcement activity and perceived safety.

For recreation, distinguish registrations, attendance, unique participants, waitlists, availability and utilization.

For youth, distinguish opportunities, applicants, participants, paid roles, volunteering, training, certifications and next steps.

For accessibility, distinguish reviews, accommodations and barriers actually removed.

For privacy and information, distinguish updates, corrections, collections, access, sharing and incidents.

For environment, define tree survival, waste diversion, energy use and other measures before claiming improvement.

For partnerships, distinguish meetings, agreements, contributions, services and resident outcomes.

Never score Saugeen Ojibway Nation as though it were a municipal department or contractor.

Measure the City's commitments and mutually agreed relationship outcomes instead.

For resident satisfaction, distinguish satisfaction, trust, confidence, process fairness and service-specific experience.

For completed and missed commitments, preserve original wording, due dates, evidence and status.

Distinguish campaign promises, mayoral commitments and formal Council commitments.

Show where another government or Council approval was required.

Do not assign failure to the wrong institution.

Set measurement and publication frequency according to the nature of the metric.

Do not create real-time dashboards merely because technology makes them possible.

Compare the same periods when making year-over-year claims.

Label partial-period and annualized figures clearly.

Keep narrative explanations separate from the underlying number.

Do not use context to hide poor performance or remove legitimate external explanations.

Use peer benchmarks cautiously and disclose why the comparator is relevant.

Do not cherry-pick peer groups to manufacture rankings.

Pair cost-per-unit and productivity measures with quality and service outcomes.

Do not treat higher productivity as permission for staff burnout.

Define uptime, downtime, cancellations, waitlists and conversion rates carefully.

Distinguish pipeline snapshots from flows over time.

Use cohort methods where survival, completion or conversion requires them.

Do not claim original backlog elimination while ignoring new backlog.

Maintain audit trails for important Scorecard calculations.

Document manual changes and corrections.

Reduce dependency on one-person spreadsheets for headline public measures.

Ensure important results can be reproduced from documented sources and formulas.

Do not depend on proprietary black-box vendor scores the City cannot explain.

Preserve historical Scorecard publications so residents can see what government reported at the time.

Avoid methodology changes immediately before elections unless genuinely necessary.

Where election-year methodology must change, publish the reason and preserve comparability where possible.

Keep the official Scorecard institutionally separate from campaigns.

Never allow Mayor or Council to suppress, recolour or rewrite unfavourable verified results.

Give residents a route to challenge definitions or data errors.

Publish methodology notes in accessible plain language with technical detail available where needed.

Structure each headline Scorecard measure around current result, baseline, target or standard, definition, source, status, trend, limitations and next action.

Treat improving data quality as useful institutional progress without confusing measurement with service delivery.

Use proportionality so measurement itself does not become a costly bureaucracy.

Retire metrics that no longer help public decisions, but publish the retirement and reason.

Never retire a metric merely because the result is politically inconvenient.

Review the metric portfolio annually for usefulness, reliability, gaming and outcome relevance.

Distinguish City-Controlled, Shared and External Context measures.

Measure Owen Sound on the parts it controls and explain material external dependencies.

Do not punish the City for weather, interest rates or provincial changes it did not control, but do measure the quality of its response where it had responsibility.

Use careful public language including Delivered, Approved, Funded, Started, Completed, Reduced, Increased, Contributed and Coincided according to their actual meanings.

Never use words such as Record, Best Ever, First, Number One, Free, Affordable, Accessible, Safe, Secure, Anonymous, Canadian or Local without a defensible definition or evidentiary basis where one is required.

Apply anti-gaming rules consistently whether the result helps or harms the government of the day.

Ask the Measurement Integrity Test before publishing every major result: Would we use exactly the same definition if the number were bad?

Use Year One to establish definitions before polishing dashboards.

Use Year Two to improve data quality and outcome measurement.

Use Year Three to identify and correct measures vulnerable to gaming or irrelevance.

Use Year Four to publish a Four-Year Measurement Integrity Audit.

Name important measurement failures as openly as measurement successes.

Hand the next Council the full metric dictionary, formulas, sources, baselines, methodology history and known data-quality gaps.

Ensure the Scorecard survives political and staff turnover.

Apply the final measurement standard to every public claim: What exactly are we measuring, compared with what, using which method, based on what source, with what confidence, and would we report it the same way if we disliked the result?

The Scorecard standard can therefore be reduced to seven rules:

Define it first.

Preserve the baseline.

Show the source.

Separate activity from outcome.

Show unknowns as unknown.

Never change the rules because the result is inconvenient.

Keep the historical record.

The purpose of measurement is not to make government look:

It is to make government:

A trustworthy Scorecard should sometimes contain:

If every measure is Green:

the public should become suspicious.

If every measure is Red:

the system may be measuring failure without managing it.

The goal is not the colour.

The goal is:

truthful information followed by competent action.

One definition. One baseline. One method. One public record. Measure the result before celebrating it, and never rewrite the measuring stick after seeing the score.

← Appendix F: Municipal Assets and Public OwnershipAppendix H: Sources, Evidence and Verification Standards →