basenull
4 min readBasenull AI Ops

"Are we doing AI right?" deserves a better answer than a demo.

Every board is asking the question. Most CIOs answer with adoption numbers and a well-rehearsed demo — which answers "are we doing AI" and says nothing about "right." Here are five operational dimensions you can actually score, and what red, yellow, and green look like on each.

AI GovernanceLeadershipReporting

At some point in the last year, your board asked some version of it: are we doing AI right? And in most companies the answer that came back was a usage number — seats deployed, tickets deflected, hours saved — plus a demo of the most impressive workflow anyone could get working the week before.

That answer addresses are we doing AI. The board asked about right — and "right" is a governance question wearing a strategy costume. It means: if this goes wrong, will it be because we were negligent? Could we withstand an audit, an incident, a regulator, a journalist?

Vibes don't answer that. A scorecard does. Here are five dimensions that are actually measurable, with a concrete test for each.

1. Inventory — can you list what's running?

The test: can your team produce, in under a day, a list of every AI system in production — models, agents, assistants, and the third-party connections (MCP servers, plugins, tool integrations) they're wired to?

Green is a maintained inventory with owners attached. Yellow is "we could assemble it in a week from configs and interviews." Red — where most organizations honestly sit — is that nobody knows how many agents exist, because connecting one requires no procurement, no ticket, and no approval.

Everything else on this scorecard depends on this row. You cannot govern what you cannot list.

2. Change control — do you find out when the surface changes?

The test: when a third-party AI dependency changes — a vendor's MCP server adds a tool, a model version is swapped, an integration widens its permissions — who is notified, and within how long?

Green: changes are detected automatically and routed to a reviewer. Yellow: you'd notice at the next scheduled review. Red: you'd notice when something breaks, or when an incident review works backwards to the change nobody saw.

3. Audit trail — can you reconstruct what an agent did?

The test: pick a real action an agent took last month. Can you reconstruct what it did, in what order, on whose authority, and whether the defined process was followed — from records, not from someone's recollection?

Green: it's a lookup. Yellow: it's possible but requires stitching together logs from several systems. Red: the systems the agent touched kept their usual logs, but the agent's own run — the steps, the ordering, the skipped approvals — left no durable trace.

4. Workforce capability — do you measure skill, or spend?

The test: for any given team, can you say where their AI capability actually is — who can delegate work to a model safely, who can verify output, who is silently not using the tools at all?

Green: capability is assessed role-by-role and re-measured; enablement money follows the gaps. Yellow: you track active usage and infer. Red: the reported metric is license count, which measures procurement, not capability — and hides exactly the teams where the investment is stuck.

5. Reporting — does the exec table see deltas, or demos?

The test: what does AI reporting to the executive team look like? Specifically: does it carry week-over-week deltas, incidents and near-misses, and cost against outcome — or highlights and screenshots?

Green: AI reporting is boring, regular, and comparable across periods. Yellow: periodic reviews with real numbers but no continuity. Red: reporting is demo-driven — which guarantees the board hears about capabilities and never about risk posture, until the day it hears about nothing else.

Scoring it honestly

Run the five tests and most organizations land red or yellow on at least four. That's not an indictment; the tooling and habits are young everywhere. What matters is what happens next, and this is where the scorecard earns its keep over the demo:

  • It converts anxiety into a work plan. "Are we doing AI right" is unanswerable; "we're red on change control, here's the plan to get to yellow by Q4" is a roadmap with an owner.
  • It survives contact with bad news. A demo-based narrative collapses at the first incident. A scorecard absorbs the incident — the row was yellow, the incident proves it, the remediation was already scheduled.
  • It's cheap. Nothing above requires a platform program. Each red-to-yellow move is weeks of focused work: build the inventory, snapshot the surfaces, define the workflows, baseline the workforce, standardize the report.

The next time the question comes — and it will come, at the next board meeting or the one after — bring five rows with honest colors and a delta from last quarter. It's a less impressive slide than the demo. It's also the only answer that actually addresses the word right, and boards can tell the difference between the executive who brought a highlight reel and the one who brought a control panel.

From the operator

Basenull AI Ops ships purpose-built tools for the IT executive whose org is already running AI in production. Governance, supply-chain security, agent ops, observability — the operational layer that usually arrives after the first incident.

Explore products