Illustrative · Kestrel Pathways is a constructed operator, not a real provider. Its shape, scale and contract mix are drawn from how the UK employability sector actually works, and every figure is indicative unless it carries a source.
Instruction 02
The balanced scorecard, on the second Tuesday
The brief for the second decision. Build a process that produces the board scorecard every month on its own, checks its own accuracy and format, and needs no email chain. Target: thirty minutes of one person time, run by anybody who has read the instructions. Today it takes two senior people a fortnight, and the board reads about the month before last.
This is a brief, not an essay, and it belongs to the fifth gate on the solution stream page: the performance regime. That gate decides what the company optimises, and this instruction is the sharpest version of it, because a board scorecard is the performance regime made visible on one page.
The ask is concrete. The board meets on the second Tuesday of every month and wants a balanced scorecard. The corporate data and reporting function, proposed as a centre of excellence, produces it twelve times a year from sources that disagree, and it currently reports the month before last, so the board is discussing a period that ended five to ten weeks earlier.
Chapter 1 · Instruction
The question
Build a process that produces the balanced scorecard for the board every month, on its own. It obtains its own data from the contributing departments and systems, computes each measure from a versioned definition, verifies its own accuracy and its own formatting, produces a version per audience, and lands before the papers go out on the second Tuesday cycle. Without an email chain, and without the last fortnight of two senior people.
The target is stated, because an instruction without one produces an improvement rather than an answer. Thirty minutes of one person time per cycle, and anybody who has read the instructions should be able to run it: no special knowledge of how the pieces fit, no relationship with the seven contributing departments, no memory of what happened last November. If it needs the director, it has failed, and if it needs the department manager, it has failed in a quieter way.
Note where that obligation sits. It is not a claim that anyone is capable of it, which would be a statement about people. It is a requirement on the instructions: having read them, a person should be able to produce the pack. Every question they still have to ask is a defect in the document, to be closed in the document, and the test below is a test of the writing rather than of the reader.
That target is what turns this from a reporting project into a process design. Twenty senior days becoming thirty minutes of anybody time is a factor of roughly three hundred, and nothing survives that except a process where the data arrives on its own, the checks run themselves, the format is fixed, and the only human act is clearing exceptions and writing a line against whatever was flagged.
The scorecard itself is part of the answer, because a process cannot be designed before the thing it produces is decided: every measure tied to a decision the board would actually take, every measure carrying a counter measure, and every measure reported at the freshest period it can honestly support.
That last clause is the second half of the instruction. The pack today reports the month before last, which means the board is looking at something five to ten weeks old, and the reason is production time rather than data availability. Operational data is complete within days of the month closing. If the process takes thirty minutes, the natural question is what can be pulled forward to the month just ended, and the answer is expected to say, measure by measure, what can and what genuinely cannot.
Two things are in scope that usually are not. The first is what to stop: the existing thirty eight page pack, and the recurring reports that will fund the capacity for this one. The second is the ownership question underneath it, because a scorecard is only as reliable as the function producing it, and that function currently reports performance and assures its own evidence.
A third fact shapes the whole answer, and it is the largest number in this brief. The pack is compiled personally by the director in charge of the department and the department manager, because it carries operational and financial information together and was judged too sensitive to delegate. It takes them the last two weeks of every month.
Price that honestly. Two senior people, half of their working time, twelve months a year: roughly two hundred and forty senior days annually, spent on assembly rather than on running the function or answering the questions the operation is waiting eleven days for. It is almost certainly the single most expensive recurring process in the corporate centre, it has never appeared as a cost anywhere, and it is the budget available to build the replacement.
It also explains the paralysis. The two people with the standing to redesign this are the two whose time it consumes, and they are consumed in exactly the fortnight when such work would have to be done.
One more absence, and it is the one that makes every improvement claim unverifiable: there is no history. The pack cannot be opened as it stood last month or as it stood the first time it was produced, so nobody can see whether a number was later restated, and no measure has a trajectory except the one drawn inside this month edition. The requirement is stronger than an archive: every prior edition has to be visible from the current one, back to the first, without anybody knowing where to look.
Then the part that consumes the second week. Every month the same thread argues the results and the formatting before anybody accepts the pack: three to five rounds of challenge, re-cut and re-order. Nobody has written down what accepted means, so in practice it means the objections have stopped.
One design principle is given rather than asked for, because it is the thing that ends the chain: a department owns its own data and does not send it. Either we agree where it parks it and how, or we agree where it already lives and the process takes what it needs automatically. Never by message.
And the mechanism underneath all of it: the whole thing runs on a single email chain with about fifteen people on it. Data goes round as attachments, drafts come back as attachments, the current version is whichever arrived most recently, and the distribution list is doing the job of an access control. Every other problem in this brief is made harder by that one fact, and it is also the cheapest to fix.
A fifth fact explains where the fortnight goes. The pack is collected rather than extracted: seven internal departments submit, each from its own system and its own timetable, and the assembly work is chasing, reconciling and adjudicating between them. Operations and finance disagree about outcomes, HR and operations disagree about headcount, and somebody senior decides in private which version the board will see.
And a fourth fact, which decides most of the design. The numbers come from four contracts commissioned by three different departments, each defining its own metrics. Much of what a board would like to see as one line does not aggregate at all, and the parts that appear to aggregate are the most misleading, because they move with the contract mix rather than with the work.
An answer that presents a list of attractive metrics without the production calendar, the counter measures and the retirement list has designed a slide rather than a scorecard, and will be scored as such.
Chapter 2 · The calendar
The cadence, which is the binding constraint
The second Tuesday falls between the 8th and the 14th, papers circulate five working days ahead, and the pack reports the month before last. Put those together and the board is discussing a period that ended between five and ten weeks earlier, having already lived through the month in between.
The important part is the fifth row. That lag is not caused by the data: operational figures are complete within days of the month closing. It is caused by the production taking a fortnight, which meant nobody could sensibly start on the day the month ended, so the cycle slipped a month and then stayed there. It is an artefact of the process, and a process that takes thirty minutes has no reason to carry it.
The honest design question is therefore three questions. What could be reported about the month just ended by the papers deadline. What is genuinely provisional at that point and should be marked as such. And what can never be current at all, which is claims validation, sustainment and audit findings, and should be dated on the page rather than quietly dropped for being old.
| Board meeting | Papers due | Period the pack reports today | How old the news is by the meeting |
|---|---|---|---|
| Tuesday 11 March | Tuesday 4 March | January | The month being discussed ended thirty nine days before the meeting. February, which everybody in the room has just lived through, is not in front of them at all. |
| Tuesday 14 April | Wednesday 7 April | February | Forty four days at the meeting, and the gap is widest exactly when the board date falls late and the reported month was short. |
| Tuesday 8 September | Monday 1 September | July | Thirty nine days, and the papers were built from data that had been sitting for a month before anybody touched it. |
| Every month | Every month | Month minus two | Between five and ten weeks old. The executive discusses a period it has already responded to, and any decision the pack prompts arrives in the quarter after the one that caused it. |
| Why the lag exists | Production, not data | The fortnight | Operational data is complete within days of month end. The month before last is reported because assembling the pack takes two weeks and nobody wanted it to start the day the month closed. The lag is a consequence of the process, not a property of the information. |
| What genuinely lags | Regardless of process | Claims, sustainment, audit | Validated claims trail by weeks, sustainment at twenty six weeks describes people who started six months ago, and audit findings arrive up to eighteen months later. Those three can never be current and should be dated on the page rather than dropped. |
derived: the second Tuesday falls between the 8th and the 14th, papers circulate five working days ahead, and the pack reports the month before last. The last column is simple subtraction and it is the most uncomfortable number in this brief.
Chapter 3 · The checks
What automatic has to mean
Automatic does not mean unattended, and it does not mean a dashboard. It means the process obtains its own inputs, computes every measure from a versioned definition, checks itself, renders itself into a fixed format, and presents a human with an exceptions list rather than a blank page and a fortnight.
The nine checks below are the ones the director and the department manager currently perform by reading. That is precisely why it takes two weeks, and why a check is occasionally missed in the last hour before the deadline. A rule does not get tired at half past six on the day the papers are due.
Note the third column. Nothing here blocks the pack. A missing input is published as a missing input, an anomaly is published as an anomaly with a line of explanation, and a failure to reconcile is published as a difference with its tolerance. The alternative, holding the pack until everything is perfect, is exactly how a cycle slips by a month and then never comes back.
| What has to be checked | How the process checks it without a person | What happens when it fails | Who clears it |
|---|---|---|---|
| Completeness | Every expected input is present for the period, at the expected grain, from every contributing source. | The measure renders as not received, naming the source and the cut off it missed. The pack still goes out. | Nobody. It is a statement of fact on the page, and the department that missed the cut off owns it in the meeting. |
| Reconciliation | The known pairs are compared automatically: operational outcomes against validated claims, HR headcount against operational headcount, employer placements against job starts. | The difference is shown as a figure with its tolerance, rather than one number being silently chosen over the other. | The data owner of the two sources, jointly, with the decision logged. It is never resolved by whoever is assembling. |
| Definition conformance | Each measure is computed from the versioned definition, and the version used is stamped on the output. | A measure computed from an out of date definition is flagged and not published until the version is confirmed. | The definition owner in the centre, as a rule rather than as a judgement. |
| Anomaly and variance | Every measure is compared with its own recent history and flagged when it moves beyond a stated threshold. | It is flagged for explanation, not blocked. A real change and a data fault look identical to a rule and different to a person. | The measure owner, who adds one line of commentary. That line is the only writing the process requires. |
| Mix versus performance | Any group level movement is decomposed into contract mix and underlying performance before it is shown. | The page shows both components rather than the net movement alone. | Nobody. It is arithmetic, and it is the check that stops the board drawing a confident conclusion that is backwards. |
| Format conformance | The document is rendered from one template: fixed layout, fixed order, page cap, every measure carrying its basis and its counter measure. | A measure without a basis or a counter measure cannot be rendered, which makes the standard structural rather than editorial. | Nobody, and that is the point. Formatting stops being a monthly negotiation because there is nothing left to negotiate. |
| Access layer | Each measure carries its sensitivity layer, and the output is produced per audience from the same source. | A restricted measure never reaches a wider audience version, because the audience version is generated rather than edited down by hand. | Nobody, monthly. The layer assignment is agreed once, in the design. |
| Restatement | Figures published for earlier periods are recomputed each cycle and compared with what was actually published at the time. | Any material difference is shown as a restatement, with the old figure, the new one and the reason, rather than the history quietly changing underneath. | The measure owner, in one line. A provisional number becoming final is normal; a final number moving is a finding. |
| Version and identity | One numbered edition at one address, with the generation timestamp, the data cut off and a generated index of every prior edition printed in the document itself. | There is no second version to be confused with, and no reader has to go looking for the last one: it is linked from this one, back to the first. | Nobody. The index is generated from the editions that exist, so it cannot fall out of date the way a maintained list would. |
summary: automatic does not mean unattended. It means the routine checking is done by the process and a human only sees what failed. The list below is what the two senior people currently do by reading, which is why it takes them a fortnight and why it is sometimes missed.
Chapter 4 · Seven contributors
Where the fortnight actually goes
The pack is not extracted from a system. It is collected from seven departments, each with its own source, its own timetable and its own account of the month, and then reconciled by hand by two senior people who are the only ones who know how the pieces are supposed to fit.
That is where the two weeks go. Very little of it is analysis. It is chasing submissions, waiting for the last one, noticing that operations and finance disagree about how many outcomes there were, deciding which version goes in, and rewriting a narrative around whichever number survived.
The fourth column is the important one. Every collision in it is a real disagreement between two departments about what happened last month, currently resolved in private by whoever is assembling the pack. That is a governance question disguised as a formatting problem, and no scorecard design survives leaving it unanswered.
| Who contributes | What they send | On what basis | Where it collides |
|---|---|---|---|
| Contract operations | Starts, outcomes, caseload, engagement, by contract and region. | The case management system, cut on the last day of the month, with late data entry still arriving for a week after. | Its outcome count and finance revenue accrual rarely match, because one counts events and the other counts validated claims. |
| Finance | Revenue, margin, cash, working capital tied up in unvalidated claims. | The ledger, closed for the month before last, which is one reason the whole pack settled on that period. | Finance sets the pace for everything, so operational measures that could be a month fresher are held back to match it, and nothing on the page says they were. |
| Employer engagement | Vacancies, placements, employer accounts, pipeline. | A CRM maintained to different standards in different regions. | Placements recorded here do not reconcile with job starts recorded in operations, and the difference has never been quantified. |
| Quality and compliance | Audit findings, complaints, safeguarding, claim rejections. | Case audits and a complaints log, both lagging by weeks. | It is the one contribution that makes the pack look worse, and it is also the one that arrives last and is most often summarised into a sentence. |
| People | Headcount, vacancies, turnover, time to hire, training completion. | The HR system, which counts posts, while operations counts people actually on the floor. | Two headcount numbers that differ by the vacancies operations is covering with overtime, which nobody reports at all. |
| Systems and IT | Availability, incidents, change backlog. | A service desk tool, reported in its own vocabulary. | Nothing collides. It is simply read by nobody, which raises the question of why it occupies two pages. |
| Bid and growth | Pipeline, re-procurement timetable, win rates. | A commercial tracker held by the growth team, some of it market sensitive. | It is the most restricted content in the pack and it sets the access level for the entire document. |
summary: the pack is not extracted, it is collected. Seven departments submit, each from its own system, on its own timetable, with its own view of what the month was like. Most of the fortnight goes on chasing and on reconciling what arrives.
Chapter 5 · The principle
You own your data, so do not send it
The answer is required to apply one principle, stated here rather than left to be discovered, because it is the thing that actually ends the email chain. A contributing department owns its data and does not send it anywhere. Either we agree where you park it and exactly how, or we agree where it already lives and the process takes what it needs from there, automatically, on the cut off.
Both modes are acceptable and most companies need both. A system of record with a decent interface is read directly. A department whose numbers live in a spreadsheet publishes that spreadsheet to one agreed location, in one agreed shape, and keeps owning it. What is not acceptable is transfer by message, because that is what creates the second copy, the version confusion, the access exposure and the fortnight.
The shift this makes is one of accountability rather than technology. Today the person assembling the pack is implicitly responsible for whether every department numbers are right, which is impossible, so they resolve disagreements by judgement under deadline. Under the principle, correctness belongs to the department that produced the data, visibly, and the centre is responsible only for taking it faithfully and saying plainly when it did not arrive.
| What changes | Today | Under the principle | Why it matters |
|---|---|---|---|
| Who holds it | Every contributor emails a copy, and the centre accumulates seven copies of the truth plus whatever it edited them into. | The department keeps its own data in its own system. Nothing is handed over, because nothing needs to be. | A copy is a fork. The moment a department sends a snapshot, two versions exist and the owner cannot correct the one that reaches the board. |
| How it moves | Attachments on a fifteen person thread, opened and re-keyed by hand. | Either the process reads the source directly, or the department publishes to an agreed place in an agreed shape. Two modes, one rule: never a message. | Transfer by message is the single mechanism producing the version confusion, the access exposure and most of the fortnight. |
| Who owns correctness | Argued in a thread, settled by whoever is most senior in it, logged nowhere. | The department that produces the data owns whether it is right, at source, before anybody reads it. | Correctness cannot be delegated to the person assembling a pack under deadline. They can only choose between numbers, which is not the same thing. |
| Shape and meaning | Whatever format each department happens to use this month. | A data contract per contributor: fields, grain, period, definition version, quality rules, owner, cut off. Agreed once, versioned after. | Without an agreed shape, automation is impossible and every month is a small integration project done by hand. |
| Timing | Chased, escalated, and waited for, which is where the second week goes. | A published cut off. The process takes what is there and publishes what is not as not received, with the source named. | Chasing is only necessary because arrival is optional. Make absence visible on the page and it stops needing a phone call. |
| Access | The distribution list, which grants everything to everybody on it. | Rights held at the source against the sensitivity layer, and audiences generated rather than edited down. | You cannot restrict what you have already emailed. Access has to be a property of the data, not of the recipients of a thread. |
| What the owner can see | Nothing. Once it is sent, the department has no idea what happened to its numbers. | Every contributor can see what was taken, when, which version of the definition was used, and how it appeared on the page. | Ownership without visibility is just blame. If a department is accountable for its data being right, it has to be able to see what was done with it. |
summary: the principle is stated by the practitioner rather than discovered by the answer. You own your data: do not send it to me. Either we agree where you park it and how, or we agree where it already lives and we take what we need from there.
Chapter 6 · Four contracts
What does not aggregate
The data arrives from four contracts, commissioned by three different departments, each with its own definitions of the events it pays for. None of those commissioners will change a definition to make a board pack tidier, and none of them should.
So the group number is the problem. An outcome rate for the company as a whole is a weighted average of three different events, and it moves whenever the contract mix moves, whether or not anything about the work changed. A board watching that line is watching the shape of the order book and calling it performance.
There is a specific trap worth naming, because it is the one that catches able people. A group rate can fall while every single contract rate rises, simply because a contract with a harder cohort grew as a share of the total. Any scorecard that shows a group trend without showing the mix alongside it will eventually produce a confident conclusion that is exactly backwards.
| The measure | Why it differs by contract | What a single group number would mean | What the board should see instead |
|---|---|---|---|
| Outcome rate | One contract pays at four weeks, another at thirteen and twenty six, the health contract counts a sustained health outcome that is not a job at all. | A weighted average of three different events, which moves when the contract mix moves even if nothing about performance changed. | Per contract, against its own trajectory, with a group line only where the definitions genuinely coincide, and the mix effect shown separately from the performance effect. |
| Sustainment | Different windows, different evidence rules, and one commissioner counts employment held with any employer while another counts the placement itself. | A number that cannot be compared with itself year on year, because the contract weights changed underneath it. | A common internal definition reported alongside the contractual ones, stated as an internal measure so nobody mistakes it for the thing the money is paid against. |
| Cost per outcome | Service fees, outcome prices and cohort difficulty differ by contract, by design. | An average that says more about which contract grew than about whether anything got cheaper to deliver. | Per contract, with the group figure shown only as a funding mix chart, never as a performance measure. |
| Caseload ratio | Priced into each bid separately, so the ratio is a commercial decision made years ago rather than an operational choice made now. | A blended ratio that describes no adviser anywhere. | Per contract against the ratio promised in its own bid, because that comparison is the one that matters and nobody currently makes it. |
| Referral volume and mix | Volumes are set by commissioners and vary with policy, seasonality and jobcentre behaviour, not by anything Kestrel does. | A group total that is read as demand for the service when it is mostly a decision taken elsewhere. | Volume shown as context rather than performance, with the assessed difficulty mix beside it, because mix explains most apparent movement in every other measure. |
| Cash and staff turnover | They do not differ. Money is money and people are people, whatever contract they sit on. | A genuine group number, and one of the few on the page. | Group level, with contract breakdown available underneath for anyone who asks. |
summary: four contracts, three commissioning departments, three theories of the same person. Each commissioner defines its own metrics and none of them will change their definition to suit a board pack, so the group view has to be built by mapping rather than by redefining.
Chapter 7 · The perspectives
What balanced has to mean here
Five perspectives, each with the question it answers for the board, the thing it must not be allowed to become, and the counter measure that keeps it honest. The general rule is in the third column of every row: a measure with no counter measure will be improved by damaging something the board is not looking at.
The participant perspective is the one most likely to be dropped for being hard to produce, and it is the one the values page commits to. If it is not on the page, the commitment is not real.
| Perspective | The question it answers for the board | What it must not become | The counter measure beside it |
|---|---|---|---|
| The participant | Is the service actually doing what we said it would, for the people we said it was for? | A satisfaction score. Satisfaction rises when a service is pleasant and undemanding, which is not the same as effective. | Effort and outcome by distance from work read together: the headline rate cannot rise while the hardest third falls without both showing. |
| The commissioner | Are we meeting the contract, and will they want us again? | A compliance dashboard. Compliance is a floor, and a board that watches only the floor learns nothing until it is breached. | Audit findings and claim rejection rate beside submission speed, so nobody can buy cash flow with defensibility. |
| Money | Are we solvent, and is the margin real or borrowed from next year? | A revenue number. Outcome revenue is a forecast dressed as a fact until the claim is validated. | Cash and working capital tied up in unvalidated claims, beside the revenue line it produced. |
| The operation | Is the machine capable, or is it being held together by effort? | A productivity metric. Productivity rises reliably when quality falls and nobody is measuring quality. | Caseload ratio and first time right beside throughput, so volume gained by thinning the service is visible in the same glance. |
| People and capability | Will we still be able to do this in a year? | A training completion percentage, which measures attendance at slides. | Adviser turnover and vacancy duration beside assessed capability, because a trained workforce that leaves is not capability. |
summary: balanced means the perspectives constrain each other. A scorecard whose measures can all improve together is not balanced, it is a single measure wearing five hats.
Chapter 8 · Baseline
What is known about the function today
Six facts about the present state, how each came about and what it costs. Every figure is indicative and stated as a baseline to be measured, not as a finding. The first act of any answer is to replace this table with the real one, because a change without a measured starting point cannot be evaluated afterwards and will be claimed as a success by whoever is still in post.
Note the board pack row. Thirty eight pages, four analyst days a month, read to page three. That is the true cost of the thing being replaced, and it is also the budget available to build the replacement.
| What it is today | The number | How it got that way | What it costs |
|---|---|---|---|
| People | 24 across group and contracts | Eleven in the group function, thirteen recruited into contracts at mobilisation and reporting to operations. | Two managers, two definitions of every metric, and nobody who can tell anyone to stop producing a report. |
| Recurring reports | About 180 | Each requested once, by somebody real, for a reason that made sense then. None has ever been retired. | Roughly three fifths of analyst time producing output whose readership has never been checked. |
| The board pack today | 38 pages, assembled by hand | It grew by accretion: every board question became a permanent slide, and no slide has ever been removed. | A board that reads the first three pages, and a pack nobody can change quickly because nobody but its authors knows how it fits together. |
| Who builds it | The director and the department manager, for the last two weeks of every month | It carries operational and financial information together, some of it before close, so it was judged too important and too sensitive to hand to anybody junior. That judgement was reasonable and was never revisited. | About twenty senior days a month, some two hundred and forty a year, which is half the working time of the two people who are supposed to be running the function. |
| History | None | Each month pack is a new document. Previous editions exist only as attachments somewhere in the email chain, and the first version is effectively gone. | The board cannot see what it was told last year, nobody can tell whether a figure was later restated, and no measure has a visible trajectory except the one drawn in this month chart. |
| Rounds of comment | Three to five, every month | The first draft goes out and the thread begins: challenges to results, requests to re-cut, and a parallel argument about ordering, wording and chart colours. | Most of the second week. The report is finished several times before it is accepted, and acceptance means the objections stopped rather than a standard was met. |
| How it travels | One email chain, about fifteen people | It started as a convenient way to gather three contributions and grew a person at a time, each added for a good reason by somebody replying to all. | The current version of the truth is the newest attachment, the audit trail is a mail thread, and the distribution list is the access control. |
| Sensitivity | Mixed, and undeclared | Operational aggregates that most of the company could see sit on the same pages as provisional financials and contract commercials that very few people should. | The most restricted item sets the access level for the whole document, so nothing in it can be reused, automated or delegated. |
| Claims evidence | About 15% of analyst time | The money arrives against evidence, so it has to be assembled, validated and defensible eighteen months later. | The one workload that cannot slip, so it wins every collision with work that would have improved the operation. |
| Questions from the operation | The remaining quarter | Answered in the order asked rather than the order that matters. | A regional lead waits eleven days, or gets a faster answer from a spreadsheet nobody validated. |
| Definitions | One document, last revised at mobilisation | Written under time pressure by people about to start delivering, and never version controlled since. | Every drift has been a reasonable reading by a reasonable person, and the cumulative effect is unknown. |
indicative: figures drawn from how functions of this size typically distribute their time, stated as a baseline to be measured rather than as a finding. The first act of any answer is to replace this table with the real one.
Chapter 9 · The catalogue
The whole range of issues underneath
The scorecard is the ask. These seventeen are the conditions it has to be produced under, and an answer that designs a beautiful page the function cannot assemble by the deadline has answered a different instruction.
Several are fixed by tooling and discipline whatever the org chart says. Two, data quality at source and the questions nobody owns, are not the data function to fix alone. An answer should say which of the seventeen it addresses, which it leaves standing, and which it refers elsewhere by name.
| The issue | How it shows up | Who feels it first | What it would take |
|---|---|---|---|
| Definitions drift | One document written at mobilisation, no version control, a decade of reasonable readings since. | The performance lead, who finds two regions counting a start differently during an audit. | Versioned definitions with a change log and an owner, and a rule that a change is an event with a date. |
| Eleven sources that disagree | Case system, commissioner portal, employer database, four spreadsheets and a scheduling tool, none reconciled. | Any analyst asked a question that crosses two of them, which is most questions. | One reconciled layer, a stated rule for which source wins, and a published reconciliation gap rather than a silent one. |
| Manual assembly | Copy, paste, adjust, send. The pack is built by hand and always has been. | Two analysts, in the last four working days of every month, permanently. | Pipelines for anything produced more than twice, and the discipline to stop hand finishing. |
| Latency against cadence | Monthly close, weekly decisions, and a board deadline that sometimes precedes the close. | The board, which discusses a provisional number as though it were settled. | A scorecard designed around what is actually knowable by the deadline, with the provisional parts marked as provisional. |
| Report sprawl | About 180 recurring outputs, 38 pages of board pack, nothing retired. | Everyone, invisibly: it is the capacity that would otherwise answer questions. | A retirement rule, somebody allowed to say no, and an annual audit of who opens what. |
| Claims crowd out everything | Evidence work cannot slip because the money depends on it, so it always wins. | The operation, told the team is in claims week. | Separating the two workloads so they stop competing for the same four days. |
| Shadow analytics | At least one unofficial spreadsheet per region, maintained by somebody whose job is something else. | The company, when two versions of the truth reach the same meeting. | Shorten the queue, then bring the shadows into the light with standards and a route into the profession. |
| Marking its own homework | The function that reports performance also assures the evidence performance is paid against. | Nobody, until an audit, which is what makes it dangerous. | Assurance signed outside the performance reporting line, with the cost in days to cash stated openly. |
| Key person risk | Two people can rebuild the claims extract. One of them wrote it. | The company, on the day one of them leaves. | Version control, documented logic, and a second pair of hands on every critical asset. |
| No history view | You cannot open the pack as it stood last month, or as it stood the first time it was produced. | Anybody asking whether something has actually improved, which is the only question a board should be asking. | Immutable numbered editions at stable addresses, plus a measure level series held separately from the document, so a trajectory survives any redesign of the page. |
| Acceptance is the absence of objection | Every month the same thread: is that number right, can we show it differently, can that chart move, before anybody says the pack is done. | The director and the manager, who do the re-cuts, and the contributors who are asked to justify their own figures again. | A written acceptance rule: what accepted means, who signs it, by when, and a frozen format so that presentation stops being reopened monthly. |
| The report lives in an inbox | Data and drafts circulate on a fifteen person email chain, with attachments as the system of record. | Everyone, quietly: nobody can say which version is current without opening three mails. | A controlled location with named access by role, version history, and a rule that the document is linked rather than attached. |
| The sensitive work stays senior | The pack is compiled by the director and the department manager because of what it contains. | The two people whose judgement the company most needs, spending it on assembly. | Layering the document by sensitivity so that most of it can be produced normally, and only the restricted layer needs a senior pair of hands. |
| Compiled by the people it judges | The department director assembles a pack that reports, among other things, the department performance and the claims its own function assured. | Nobody, until an auditor or a non executive asks who checked it. | A named reviewer outside the line for the sections that judge the function, which costs a day and removes the objection permanently. |
| Data quality at source | Mandatory fields completed under time pressure, so they are complete rather than true. | Every analysis built on them, silently. | Fewer mandatory fields designed with advisers, validation at entry, and a feedback loop showing them what their data produced. |
| A different return per commissioner | Four contracts, four formats, four cuts of the same underlying facts. | The analysts, at every mobilisation, rebuilding from nothing. | A reusable internal model with per contract presentation on top. |
| Nobody owns the harder questions | Sustainment beyond twenty six weeks, effort per participant, repeat referral rate: nobody job to produce. | The participants the company said it would not park, and the values page. | Naming them as standing products of the function, on a schedule, whether or not anybody asks. |
summary: the scorecard is the ask, and these are the conditions it has to be produced under. An answer that designs a beautiful scorecard the function cannot produce by the deadline has answered a different instruction.
Chapter 10 · Goals
What the department is for
A function cannot be designed before somebody says what it is for, and a scorecard is an expression of a purpose whether or not the purpose was stated. These five goals are proposed here so the design has something to serve.
Each states how it is measured, and three of the five measures cannot currently be produced at all. That is the most useful finding on this page: a department that cannot measure its own goals is in exactly the position of the operation it reports on.
An answer may argue with any of these goals. What it may not do is propose a scorecard without saying which purpose it serves.
-
Every number means one thing
Goal 01- What it means
- A metric has one definition across the company, versioned, owned, with a change log. Where a contract needs a different cut, it is a presentation of the same definition rather than a second meaning.
- Measured by
- Proportion of reported metrics carrying a current versioned definition, and the count of unlogged definition changes in the period. The second should be zero and currently cannot be produced at all.
-
An answer while the decision is still open
Goal 02- What it means
- The function exists to change decisions. An answer arriving after the decision was taken has produced nothing, and the asker learns to stop asking.
- Measured by
- Elapsed days from question to an answer the asker acts on, as a distribution, and the proportion delivered before the decision date the asker stated. Nobody currently records that date.
-
The board reads a true page, not a tidy one
Goal 03- What it means
- What goes to the board is the position as it actually is, including the parts that are provisional, missing or unflattering. A scorecard that is never uncomfortable is not measuring anything that matters.
- Measured by
- Proportion of scorecard measures carrying a stated basis and a counter measure, and the number of board items where the number was later found not to support the conclusion drawn from it.
-
Claims that survive an audit eighteen months later
Goal 04- What it means
- The money arrives against evidence, so the evidence is the product. Complete, traceable to source, defensible by somebody who was not in the room when the outcome was claimed.
- Measured by
- Claim rejection rate, audit findings per thousand claims, and days from outcome to submission, read together because any one can be improved by damaging another.
-
The operation can answer most of its own questions
Goal 05- What it means
- A centre that hoards skill makes itself permanent and the operation dependent. The test of this function is what a region can do correctly without it.
- Measured by
- Share of questions resolved without a centre analyst, alongside an assessed test of analytical literacy rather than a self declared one, twelve months after any change.
Chapter 11 · Comparators
The boundary of the comparison
Five production models will be scored, and they were named before any weighting was set. One: keep the present arrangement, email chain included, and simply shorten the pack. Two: structured submission, where every department files to a template in one controlled location by a published cut off and the centre assembles by hand from what is filed, with no email in the loop. Three: automated pull, where a reconciled data layer takes what it can from source systems and contributors supply only what genuinely cannot be pulled, which is commentary. Four: full automation with exception handling, where everything computable is computed, the nine checks run themselves, and humans only clear exceptions and write one line each against flagged measures. Five: buy a reporting platform and rebuild the pack inside it.
The set contains carrying on, because a brief that excludes the status quo has already decided, and it contains at least one option somebody credible would argue for, which is the fifth. Two options were considered and excluded, and are named rather than omitted: outsourcing production, excluded because assembling it is where the function learns what is true; and asking the board to move its date, excluded because the cadence is fixed by the group calendar and the instruction is to work inside it.
A note on scope, since it decides the scoring. Options three and four require the reconciled data layer, the versioned definitions and the access layering to exist. Those are not free and any answer proposing them must include the cost of building them rather than assuming them into place.
Chapter 12 · Criteria
How the answer will be graded
Five criteria, weighted, with the method of scoring named for each. The threshold is sixty out of a hundred: below that an option is not recommended regardless of comparison, because continuing honestly with the current pack is always available.
Criterion three carries a floor in addition to its weight. A design that leaves the operational measures reporting the month before last scores below half on it and is not recommended, whatever it earns elsewhere. Automating a fortnight of assembly and keeping the five week lag would be an expensive way to preserve the problem.
| Criterion | Weight | Why it carries that weight | How it is scored |
|---|---|---|---|
| Thirty minutes, from the instructions alone | 30% | The target is the instruction: one person, thirty minutes, working from what is written. It is also the only version of this that ends the single point of failure and returns two hundred and forty senior days a year. | Measured, not estimated: a person from outside the department, given the written instructions and nothing else, produces a full cycle while somebody times it. Scored on the clock and on whether they needed to ask anybody a question. |
| Checks itself | 25% | Accuracy currently depends on two people reading carefully under deadline. A process that produces faster without checking harder is worse than what exists now. | How many of the nine checks run automatically, what each does on failure, and whether a measure can reach the page without a basis, a counter measure and a definition version. |
| Closes the reporting lag | 20% | The board currently discusses the month before last, five to ten weeks after it ended, because production takes a fortnight rather than because the data is unavailable. A thirty minute process that still reports month minus two has automated the wrong thing. | For each measure, the earliest period it can honestly support at the papers deadline, and the lag actually achieved. Floor: a design that leaves the operational measures at month minus two scores below half here and is not recommended, whatever it earns elsewhere. |
| Ends the email chain | 15% | The chain is the version control, the access control and the audit trail, and it is none of those things. It also carries operational, financial and commercially sensitive material to fifteen people at one access level. | Whether the document is generated per audience from one source at one address, and whether anything at all is still attached to a message. |
| Cost, and what it retires | 10% | Headcount is flat, so the build is funded by stopping things. An answer that needs money the company has not got is a wish with a Gantt chart. | One off build cost and steady state cost as ranges with their basis, and the named list of reports and pages retired to pay for it. |
summary: weights fixed and published before any option was scored, and the comparator set named before the weights. The instruction is a process rather than a page, so the weights sit on production: whether it runs without hands, whether it catches its own errors, and whether it can do both inside a deadline that sometimes precedes the close.
Chapter 13 · Fixed
The constraints that are not negotiable
The board date does not move. Second Tuesday, papers five working days ahead, including in the months where that lands on the 1st.
Headcount is flat this financial year. The replacement is funded by retiring existing outputs, not by hiring, and the answer names what is retired.
Commissioner reporting is unaffected. Contractual returns keep going out in their required formats, on time, throughout any transition.
No identifiable participant data goes into board papers. Aggregate only, with a minimum cell size stated in the design.
Every number on the page carries its basis: sourced, derived, provisional or indicative. A figure with no basis does not go on the page, whoever asks for it.
The written instructions are the deliverable, not a by product. They must be sufficient on their own: a person who has read them should be able to run the cycle without asking anybody, and they are kept with the process rather than in somebody head or somebody inbox. Where a reader still has to ask, the instructions are amended, not the person.
The format is frozen for twelve months. Presentation is agreed once, in the design, and is not reopened in a production cycle. A request to change the layout is a change request for next year edition, not a reason to re-cut this month.
A dispute about whether a number is correct is resolved by the definition and the reconciliation rule, logged, and it does not delay the pack. A dispute about what a number means belongs in the board meeting, which is the forum that exists for it.
Accepted has a written meaning: named signatory, stated checks, and a time. The pack is accepted when those are satisfied, not when the thread goes quiet.
Every edition is immutable and addressable, and every prior edition is reachable from the current one. The version index sits in the document itself, listing each earlier edition back to the first, including any reissue inside a month, so a reader never has to know where an archive lives. The measure level series is held outside the document, so a redesign of the page never breaks a trajectory.
The document stops travelling as an attachment on a standing email chain. It lives in one controlled location, access is by named role against the sensitivity layer, versions are numbered, and people are sent a link rather than a file. Participant level data is never attached to anything.
No data arrives as a message or an attachment. Each contributor either exposes its source for direct reading or publishes to the agreed location in the agreed shape, and the centre never holds a private copy that the owner cannot correct.
Every contributor can see what was taken from it, when, under which definition version, and how it appeared. Accountability for correctness sits with the owner, and it is only fair if the owner can see the consequence.
One cut off for every contributing department, published in advance, the same every month. What has not arrived by the cut off is shown on the page as not arrived, with the department named, rather than chased into the evening by a director.
Commissioner definitions are fixed. The group view is built by mapping contractual measures onto an internal model, never by asking a commissioner to count differently or by quietly redefining a contractual term internally.
Sensitivity is handled by layering, not by restricting the whole document. Operational aggregates, provisional financials and contract commercials have different audiences, and the design states which layer each measure sits in and who may see it.
The sections that judge the data function are reviewed by somebody outside its line before they go to the board. The director may write the pack; they may not be the only person who has checked the part that scores their own department.
And no definition may be changed in the run up to a board to make a measure land better. That is a constraint rather than a criterion, so it is not tradeable against anything.
Chapter 14 · Risk
The risks this decision has to address
Five mechanisms, each with how it works, who carries it and what it would take to measure. The first is the one that matters most: the scorecard becomes the operating model, so whatever is left off the page is left out of the company, and the thing most likely to be left off is the thing hardest to produce by the 1st.
-
The scorecard becomes the operating model
Unquantified- Mechanism
- What the board watches is what the executive manages, and what the executive manages is what two thousand people do. Within two quarters the organisation becomes whatever is on this page, including the parts that were chosen because they were easy to produce.
- Who carries it
- Whoever is served by the things that did not make the page, which on current evidence is the third of the caseload furthest from work.
- To measure it
- Compare what the scorecard measures with what the values page commits to. Any commitment with no measure is a commitment nobody is accountable for.
-
Provisional numbers harden on contact with a board
Unquantified- Mechanism
- A figure marked provisional in a pack is quoted as settled in the meeting, then quoted again externally. The caveat does not travel with the number.
- Who carries it
- The executive who repeats it to a commissioner, and the credibility of every later number.
- To measure it
- Track restatements: how often a figure presented to the board was materially different when final. If nobody is tracking that, the board does not know how much to trust what it reads.
-
The monthly argument decides the number
Unquantified- Mechanism
- A figure is challenged, re-cut, challenged again, and the version that survives is the one nobody senior objected to. Two entirely different things are being settled in the same thread: whether the number is correct, which is a matter of fact, and whether people like what it shows, which is not. Under deadline they resolve together.
- Who carries it
- Whichever department is least senior in the thread, and the board, which receives a negotiated figure believing it is a measured one.
- To measure it
- Rounds of comment per cycle, and a log of every change made after first draft marked as either a correction or a presentation change. If corrections are rare and presentation changes are many, the argument was never about accuracy.
-
The past is whatever this month pack says it was
Unquantified- Mechanism
- With no archive, a comparative figure is recomputed each month from current data and current definitions. Numbers move for good reasons, nobody can see that they moved, and the company loses the ability to answer whether anything has actually improved since it started asking.
- Who carries it
- The board, which cannot audit its own minutes against the pack that produced them, and every improvement claim the company makes to a commissioner.
- To measure it
- Whether an edition can be reopened exactly as issued, and whether restatements between editions are published. Neither is currently possible, so the honest answer today is that the trajectory is unverifiable.
-
The newest attachment is the truth
Unquantified- Mechanism
- When a document is assembled by email, the current version is whichever file was sent most recently, and there is no way to tell from the outside. Somebody eventually presents last week figures with this week cover, and nobody notices until a number is questioned.
- Who carries it
- The board, and the two people who will be asked how it happened.
- To measure it
- Count versions in circulation for one cycle. If the answer is more than one, the risk is live, and it is currently certain to be more than one.
-
The distribution list is the access control
Unquantified- Mechanism
- A chain grows a person at a time, each addition reasonable at the moment it happens. Within a year, fifteen people receive operational, financial and commercially sensitive material together, because the pack has only one access level and the chain has only one reply all.
- Who carries it
- The company, in confidentiality and in data protection terms, and whoever is on the list who should have seen one layer and received all of them.
- To measure it
- List the recipients and the layers each of them needs. The gap between the two is the exposure, it takes an hour to produce, and nobody has produced it.
-
The agreed location becomes another inbox
Unquantified- Mechanism
- Departments park files as asked, nothing validates them on arrival, and within two quarters the folder holds late, malformed and duplicate submissions that somebody has to read. The email chain has been recreated with worse search.
- Who carries it
- The centre, which is back to manual reconciliation, and the deadline.
- To measure it
- Quality rules run at the moment of arrival, with the result visible to the contributor immediately. Count submissions accepted first time: if that rate is not near total, the contract is wrong or the rules are not running.
-
Reading a source nobody told you had changed
Unquantified- Mechanism
- Direct reading creates a dependency the owning department cannot see. They rename a field, change a code list or fix their own data for perfectly good reasons, and the scorecard breaks or, worse, quietly reports something different.
- Who carries it
- The board, if it is subtle enough not to break anything visibly.
- To measure it
- Schema checks on every read, versioned contracts, and a notification obligation both ways. The count that matters is silent changes detected after the fact rather than before.
-
The number is negotiated rather than measured
Unquantified- Mechanism
- Two departments submit figures that disagree. Somebody assembling the pack under deadline chooses one, or splits the difference, or takes the one with the better story. It is done in good faith, it is never recorded, and over time the pack becomes a consensus rather than a measurement.
- Who carries it
- The board, which believes it is reading data, and whichever department is quietly overruled each month.
- To measure it
- Log every reconciliation: the two figures, the one chosen, and why. If the log is empty, either nothing ever disagrees or nobody is writing it down, and only one of those is plausible.
-
The group line that is really a mix chart
Unquantified- Mechanism
- Aggregating measures defined differently by four commissioners produces a number that moves when the contract mix moves. A group rate can fall while every contract in it improves, and the board takes a decision about performance on the strength of a change in the order book.
- Who carries it
- Whichever contract is blamed for a movement it did not cause, and the credibility of the pack when somebody eventually notices.
- To measure it
- Decompose every group movement into mix and performance, monthly. It is one extra calculation and almost nobody does it.
-
Green everywhere, quietly
Unquantified- Mechanism
- Targets get set where they can be met, thresholds drift, and a scorecard that is green every month for a year is describing its own calibration rather than the business.
- Who carries it
- The board, which stops reading, and then the company, which loses its early warning.
- To measure it
- Proportion of measures green over rolling twelve months. Above about eighty percent, the targets are the problem rather than the performance.
-
Sensitivity is the reason nothing ever industrialises
Unquantified- Mechanism
- Because the document is important and mixed in sensitivity, it is done by hand by senior people. Because it takes them a fortnight, there is no fortnight left in which to fix it. Because it is never modelled, versioned or automated, it stays a job for senior people. The loop is stable, self funding and can run for a decade, and every participant in it is behaving sensibly.
- Who carries it
- The director and the department manager, who lose half their working lives to it, and the company, which has a single point of failure on its board agenda and a leadership team with no capacity to improve anything.
- To measure it
- Senior days per cycle spent on assembly rather than interpretation, tracked monthly, and a simple test: could next month pack be produced if both were unavailable.
-
It grows back to thirty eight pages
Unquantified- Mechanism
- Every board question becomes a permanent addition. Nothing is ever retired, because retiring a measure requires somebody to say that a director question was not worth a recurring page.
- Who carries it
- The function, and every question the operation could have had answered with that capacity.
- To measure it
- A hard cap agreed in advance, with a rule that a new measure displaces an existing one. Count the page length every month, which is the cheapest audit available.
-
Managing a period everybody has already left
Unquantified- Mechanism
- A board discussing the month before last is responding to something the operation noticed weeks ago and has already acted on. The discussion produces instructions about a situation that has moved, and the executive learns to treat the pack as ceremony rather than as information.
- Who carries it
- The board, whose questions arrive a quarter late, and the operation, which receives direction about a problem it already fixed.
- To measure it
- For each board action, the date of the data that prompted it against the date the operation had already acted. If the second reliably precedes the first, the pack is not informing anything.
-
The lagging truth never arrives
Unquantified- Mechanism
- Sustainment and audit findings lag by months. A scorecard built for timeliness quietly drops them, and the board ends up watching only the fast numbers, which are the gameable ones.
- Who carries it
- Everybody, eventually, because the slow numbers are the ones that say whether any of it was real.
- To measure it
- Whether the lagging measures are on the page at all, dated honestly as describing an earlier cohort rather than omitted for being old.
Chapter 15 · The test
What a complete answer has to contain
The actual scorecard, drawn, at the size it would be read. Not a description of one. For every measure: its definition, denominator, exclusions, source, basis, owner, the decision it informs, the threshold at which the board would act, and the counter measure sitting beside it.
The process that produces it, end to end: what is pulled and from where, what is filed and by whom, the cut offs, the nine checks and what each does on failure, how the exceptions list reaches a human, and what that human is actually required to do. If the answer is more than an hour of work a month for anybody, say why.
A production calendar walked against a real month, showing for each measure the period it reports, whether it is final or provisional at the papers deadline, and what appears on the page to say so. Including the months where the papers are due on the 1st, which is where any pull forward will be hardest.
For every measure, a ruling on aggregation: does it combine across the four contracts, and if it does, what the combined figure means and how mix is separated from performance. Measures that do not aggregate are shown per contract or not at all, and that decision is stated rather than implied by the layout.
A retirement list. Which of the thirty eight pages and which recurring reports stop, who is told, and what capacity that releases.
An acceptance rule, in writing: what has to be true for the pack to be accepted, who signs, by when, and what happens to a late objection. Plus the dispute path that separates a correction from a disagreement about what the result means, so the second stops delaying the first.
A history model, navigable from the current edition rather than filed somewhere else: the version index in the document, every prior edition openable exactly as issued back to the first, the measure trajectory held independently of the page, and restatements shown when a published figure later changes. The first edition produced under the new process is also the baseline, and the answer should say so explicitly.
A transport and version model: where the document lives, how contributions arrive without an email chain, who may open which layer, how a version is identified, and what the board is sent. If it is still an attachment to fifteen people at the end of the design, the design has not finished.
A data contract for each of the seven contributors, written out: read directly or published, the exact location and shape, fields, grain, period, definition version, quality rules applied on arrival, the cut off, the named owner, and what the process does when the contract is not met. This is the artefact that replaces the email chain, and an answer without it has described the principle rather than applied it.
A collection model: for each measure, the contributing department, the named owner in it, the source system, the cut off, and the rule for what happens when two departments disagree, including who adjudicates and where that is recorded.
A production and access model: which layer each measure belongs to, who may see it, who assembles it, and what genuinely still requires the director and the department manager. State the senior days returned, because at twenty a month that number is the business case and it is larger than anything else in this instruction.
A rule for change: how a measure gets added, what it displaces, and who is allowed to refuse a director request for a permanent new page.
A position on each of the seventeen issues in the catalogue, and a statement of which of the five departmental goals the design serves, with any it cannot serve identified rather than quietly dropped.
The written instructions, at the level of detail somebody unfamiliar can follow, plus the timed proof: a named person who has never produced this pack, running one cycle from the instructions alone, with the elapsed time recorded and every question they had to ask listed. Each question is a defect in the instructions, and the answer says how each was closed.
A dry run: the scorecard produced for the month just gone, with real numbers, before anybody is asked to approve the design. Including the parts that will embarrass somebody, because a scorecard that is comfortable in its first month was designed to be.
And a named owner at a level that holds both sight and control. On the actor model that is group, and this is one of the few decisions where group can genuinely see what it is deciding about.
Instruction 02 is open
The brief is fixed: cadence stated, comparators named, weights published, floor declared. The answer will be posted here and graded against exactly these criteria. If you have built one of these for a board, the most useful thing you could tell me is which measure you regret putting on it.
Tell me what is missing