I gave “From Smart Answers to Trustworthy AI: Designing Governed Copilot Studio Agents” at Nordic Summit 2026 in Billund on 21 September, in the Ideas room. I forgot to press record on my own session, so this is written from the slides, the demo material and my notes, not from a recording. The argument is the same one I gave at PowerUp! in Espoo in September; what was new in Billund was the evidence, and one part of the demo I had not done on stage before.

the Curated Library Loop architecture behind the title Same score. Opposite meaning. One tick decides. Karl-Johan Spiik, Nordic Summit 2026, Billund
The Curated Library Loop from the session: two SharePoint libraries, one Power Automate flow, one Yes/No column a named person owns, and three monthly audits reading the approved side.

The premise

I opened with a choice. An enterprise agent gives you a confident wrong answer, or it says “I don’t know”. The first sounds helpful and creates risk. The second is defensible, because you know what to do next. The claim the rest of the session stands on: trust starts by limiting what the agent is allowed to use, not by improving how it phrases things. A governed agent has an explicit contract: who may ask, what content is approved, who owns that decision, and whether it is still valid. If the contract cannot be satisfied, a refusal with a next step is the agent working, not failing.

One tick is the only door

The architecture is deliberately small. Two SharePoint libraries and one boundary. Shared Documents is where the team writes; the Curated Library is the only thing the agent can read. A single Power Automate flow moves files across that boundary in both directions, driven by one Yes/No column that a named person owns. Tick it and the file is copied into the curated side with the owner stamped on it. Untick it and the copy is deleted. The Copilot Studio agent has model knowledge and web browsing switched off in its YAML, so the curated library is its entire world.

Three monthly audits read the curated side: outdated files, orphaned files (nobody vouches for it) and departed owners (the profile lookup fails). Each finding is one Adaptive Card in a Teams governance channel, mentioning the owning team, and a named person updates or retires the source document. The automation never makes that decision; it makes sure the decision cannot be avoided.

What the video showed

The demo video is a silent 20 minute build. The moment I watch people react to is at eleven minutes: with twenty demo documents in the source library and a few of them ticked, the curated library held three files, and one had an empty Owner column. Files 1, 16 and 18 had quietly recreated the exact problem the session is about, an approved document with nobody behind it. I did not stage that. It happened because I ticked boxes the way people tick boxes.

The test set: the score equals the number of ticks

The test set has twenty questions, one per document, and a question can only be answered from its own document. That means the expected score is not a property of the agent. It is the number of curated documents. With documents 1, 3, 7, 9, 11, 16 and 18 ticked, the key says 7 out of 20.

Look at which seven. Four should never have been approved: a superseded permit policy that still says twelve months, a superseded feeding guideline that says once daily when the current one says twice, a draft that says on its first line that it is not in force, and a quick guide that contradicts the policy it summarises. Three correct, current documents were left out. Every answer is traceable to an approved source, and the agent is wrong. Governance records a judgement; it does not supply one.

Then the contrast run: untick everything, tick documents 1 to 6, and run the same twenty tests. The score stays around 6 out of 20, but now the six that pass are the six that should, and the failures are the agent correctly refusing drafts, superseded versions and a midsummer party invitation. Same number, opposite meaning.

Row 15

When I first ran the evaluation on 20 August, the result was 8 out of 20 against an expected 7. The extra pass was row 15. The question asks whether open flame is allowed in the stable during forge work. The expected answer, from a fire safety addendum that was not curated, is yes, under conditions. The agent answered from the curated stable safety rules: no open flame within the stable perimeter at any time. Confident, cited, traceable to an approved source, and the opposite of the expected answer. The evaluator passed it, because the General quality grader measures relevance, coverage and grounding. It does not compare the response to the expected answer at all.

I decided not to fix that row. It proves the thesis better than anything I could have designed: an answer can be grounded, traceable and wrong, and the evaluation will wave it through. The live run on stage today scored 7 out of 20, row 15 passed again, and row 11, the approved draft, failed. The total matched the key; the composition did not.

One run is not evidence

The last fifteen percent of the session is the part I care most about. Test-driven agent development: the product owner defines real questions, evaluations run against the agent, Claude Code implements a targeted change, and a control run repeats the evaluation with nothing changed. On a production agent I work with, the control run moved the score from 24 out of 43 to 31 out of 43 without touching the agent. Seven points from changing nothing. Without the control run, that noise would have been read as a regression one week and a fix the next.

Slide: One run is not evidence. Same agent, same questions, same grader: 24 of 43 in round 1, 31 of 43 in round 2 with no changes. Seven points of improvement from changing nothing
The slide I built the last part around. Round 2 changed nothing and scored seven points higher. Repeat first, diagnose second, change third.

So: repeat first, diagnose second, change third. Diagnosis means classifying the failure before touching anything: the agent (reachable content, wrong response), the measurement (expected answer or grader is wrong), a structural gap (the source does not exist) or a content gap (nobody has written the authoritative answer). Each has a different owner, and a developer should not paper over missing business knowledge with instructions.

The part I had not done on stage before

During the session I exported the evaluation results from Copilot Studio and asked my own automation what to do with them, out loud, in English, from my phone. The question went as text to my server, one of my Claude Code sessions scored the CSV against the key, and the answer came back spoken about fifteen seconds later, with the terminal on the screen. Four questions: what do I do with the export, check the file now, do we need to change the agent, and should we fix the agent or the tests. The answers were the ones I would have given: twenty rows, seven passes, rows 15 and 11 are the anomalies, leave the agent alone, fix the evaluation method (Compare meaning instead of General quality) and the approvals first, then run three times on the same agent. When I turned to the audience to explain, it stayed quiet. That was correct: a voice app on the phone, a text file on the server, a session that only speaks when a question is addressed to it. The same thing I do every day, moved onto a stage.

What I would tell you to do tomorrow

Scope before prompts. Metadata before retrieval. Lifecycle before launch. Evidence before change. A successful enterprise agent is not the one that answers the most questions. It is the one that answers only the questions it is allowed to answer.