Boring Magic

@boringmagi.cc.web.brid.gy

Product and design consultancy. Helping you build products and services that meet people’s expectations. No hype, no fluff. Good, straightforward stuff that just works.

Rapid AI prototyping at Sport England

**In March 2026, Boring Magic joined Register Dynamics and Oxford Insights inside Sport England’s Data & AI Lab to explore what AI could actually do for the organisation. The brief was vague. The constraints were real. Five weeks later, there were working prototypes in staff hands.** We used discovery and rapid prototyping to produce credible, testable prototypes, not abstract recommendations. > “The initial brief was quite techy in nature: do something with data and AI. Steve’s thinking helped focus our work on solving an actual problem that the business faced.” Simon Worthington, Register Dynamics ## The context Sport England is not a startup. It’s a public body with a mission to help people across England be active, and with that comes real governance. Sport England’s AI position statement, published in February 2026, is clear: AI can support tasks like processing applications, analysing monitoring data, and summarising reconciliations. But it cannot make decisions. That’s the right guardrail, one we used to define the design space. The organisation was dealing with slow, duplicative reporting, clunky legacy systems, and staff who were understandably sceptical about whether AI would actually help. The Data & AI Lab existed to test what was genuinely achievable, and to build trust with staff through working prototypes rather than promises. Our job was to find the operational friction worth solving, and show what ‘useful’ could look like within those constraints. ## What we did We used Lean UX and rapid prototyping to move through the problem space and solution space at the same time. The approach was structured around a few clear principles. 1. **Design with staff, not just for them.** We used participatory methods so that Sport England staff shaped the prototypes, not just reviewed them. The objectives-setting prototype, for example, was heavily influenced by the Partnerships team’s own ideas. 2. **Map constraints before building.** We reviewed Sport England’s AI policy and the government’s AI Playbook before writing a line of code or designing a single screen. That’s how we ended up working inside Copilot rather than proposing a new system. 3. **Augment, don’t replace.** In two prototypes, human review was the primary pattern. AI-assisted analysis was a secondary option users could choose. Staff stayed in control of review stages and the outputs. 4. **Solve for now, point to the near future.** Quick wins grounded in real operational problems, not multi-year transformation programmes. We ran regular show-and-tells to keep feedback loops open even when staff availability was limited. ## What got shipped in five weeks Two areas of operational friction became the focus. Both produced working prototypes. * **Reviewing grant applications.** The Movement Fund receives around 300 grant applications per week. Decisions are made manually, and the 12-week target is regularly stretched to 10 weeks in practice. We spoke with the team responsible for investment management and mapped the assessment pipeline: submission, assessment, recommendation, peer review, decision. The prototype used AI to pull key information out of applications to support a recommendation, with the goal of making peer review meetings more focused and reducing the number of contentious cases that needed extended discussion. Human review remained the default throughout. * **Analysing partner reports and setting objectives with Active Partnerships.** This became three distinct approaches: a set of prompts to help staff analyse reports inside Copilot without hitting context window limits; an automated workflow to reduce the friction of Copilot’s chat-only interface; and a prototype UI for setting objectives with partners collaboratively, based on the analysis. Copilot was the deliberate choice throughout. It was the approved, accessible tool already inside Sport England’s ecosystem. Using it meant the prototypes were immediately feasible within existing governance, rather than dependent on new procurement or IT approval. That’s not a compromise. Working within existing constraints helped us to move fast and deliver value early. ## What made the work credible Like many public sector organisations, the constraints were real – and worth being honest about. * With only 1 to 2 hours with staff for context-gathering, we had to be disciplined about what we explored and how we used their time. * We lost one week of prototyping time to delays getting access to accounts. A slower start, but two Lab accounts now exist inside Sport England for future work. * One prototype was paused after speaking with the relevant team. But that was a positive: the work was governed by fit and usefulness, not by a need to ship everything. None of that derailed the project. We were still able to keep moving forward. The show-and-tell rhythm meant feedback kept flowing even when staff couldn’t respond to messages. The scrappy-first approach meant we weren’t precious about discarding ideas that didn’t hold up. And being honest with Sport England teams about what we were doing, and what we weren’t, built the kind of trust that makes future work easier. The teams we worked with gave us some greaet feedback: > “If you’d have asked us to spend 4 hours in a workshop at the start of this, we’d have said we didn’t have the time. But seeing what you’ve done now, we’d find 2–3 days at least to spend with you.” That’s the model for rapid discovery and prototyping: practical progress under real conditions, not a polished case study built under ideal circumstances. ## Why this matters Most public-sector and government-adjacent organisations aren’t short of AI ambition. They’re short of working methods that suit an ambiguous brief, or work within policy and procurement constraints, keep staff in control, and still get to something useful and testable quickly. That’s the gap we filled. Not a big, multi-year transformation, and not a strategy document for someone else to act on. A four-to-eight week discovery and prototyping project that finds what’s valuable, viable, feasible and usable, and leaves behind something the team can actually build on. If you’re working through a similar brief, or trying to figure out where AI can genuinely help your organisation without the hype, get in touch.

boringmagi.cc

Reality drift

You’re driving down the motorway. You’re halfway through the journey, with a few hours left till you arrive at your destination. It’s been an okay drive so far. Sure, it’s long and has taken up most of the waking day, but the route and plan has paid off so far. Suddenly, a light starts flashing on your dashboard. The car is running out of fuel. _How can this be?_ The fuel gauge says the tank is half-full. You indicate and pull over into the hard shoulder as the vehicle slows to a halt. You try starting the car. It chokes the familiar hoarse chatter, but never turns into the comforting rumble. You look at the fuel gauge again. Half-full? _How did I get here?_ This is reality drift: the gap between what you believe or assume is true and what actually exists or happened in the real world. The information supporting your plan was wrong, and now you’re in a place you didn’t intend to be. In adverts, we’re told that large language models are capable of powerful things – making sense of huge datasets, surfacing insights we couldn’t find as rapidly ourselves. Some of that is true. But the performance showcased in AI marketing is a fair distance from the experience on the ground. Over the past year or so, I’ve worked with teams who have tried to use general AI – not specialist tools – to summarise or analyse large blocks of relatively unstructured text. It’s one of the most common AI use cases right now, and one of the riskiest if you’re not careful to check that the output is truthful. ## Analysing a year’s worth of reports A team working with a large network of nationwide partners used Copilot to analyse a substantial spreadsheet and produce a report: including an executive summary, trends over time, a heatmap, and a narrative summary of each partner. The aim was to use the report to set objectives with the partners, goals for the next year. After inspecting the report, the team felt that something was off. When we looked at it together, we found that only a handful of partners had been mentioned in the trends. More worryingly, some of the narrative summaries described partners as performing well against certain outcomes – even though there was no evidence in the data to support those claims. The report looked credible. It had structure, narrative, and numbers. It just wasn’t accurate. ## Synthesising user research A user researcher with dozens of interview transcripts, each lasting over 40 minutes, had analysis to perform. They turned to the chatbot provided by their organisation’s enterprise software. They loaded it with all the documents and asked it to identify themes, group quotes, and write up findings. A strong theme running through the reports was that users struggle to find the right page they need. Navigation wasn’t quite right. So they recommended a card-sorting exercise to rethink the information architecture. Three weeks into the project, someone else on the team scanned the interview transcripts. A few users had mentioned a workaround that suited them, which in itself was a solution to help simplify the information architecture. They implemented the change quickly and tested it, seeing in the analytics that fewer users dropped off or got lost. A simple change lurking in the raw transcripts that AI missed…and the team nearly missed it too. ## Why does this happen? These aren’t random failures or bugs that will be ironed out. This is due to how large language models are built and structured. ### Context window limits Every LLM has a context window, a maximum amount of text it can hold and process at one time. Think of it as the model’s short-term memory. Many AI products are configured to still produce a response even when the input exceeds what the model can handle. It doesn’t always tell you it’s working with incomplete information. It just carries on and produces an output. ### Hallucinations LLMs are statistical models trained on vast amounts of text. They learn to predict the next likely word or token based on what came before – it’s not based on ‘truth’ or ‘facts’. When the model lacks sufficient context, it fills the gaps with plausible-sounding content based on patterns in its training data. The result can look authoritative while being wrong, a fabrication. ### Non-determinism Unlike traditional rule-based systems, LLMs operate probabilistically. Give the same prompt twice and you’ll often get two different responses. This makes it hard to validate outputs against each other, and hard to trust any single run of an analysis as definitive. ## Drifting, drifting Reality drift doesn’t happen to everyone who uses AI. In fact, plenty of leaders have made decisions far from reality without even using AI. But that’s another problem entirely. People are particularly prone when using off-the-shelf AI to produce insights without understanding the inherent risks. There are companies developing specialist AI tools who take care over accuracy, reliability and trustworthy outputs. But most enterprise LLMs, like Copilot, are generic – they can’t guide you, and they don’t let you configure many settings. We’re not saying enterprise tools are off the table, particularly not for teams working in the public sector, where funding is tight. You’ll just need to be aware of a few things. ## What to do to mitigate reality drift Reality drift is a risk you can manage, but not eliminate entirely. Here’s what we’ve found helps. ### Treat all AI-generated summaries with professional scepticism Review claims, especially narratives and evaluative statements. Check them against source documents before acting on them. ### Use an evidence-first approach Rather than asking AI to produce a polished summary in one go, extract the raw evidence first – direct quotes, data points, observations, references – and validate those against source material. Classify and organise the evidence yourself, then use AI to help with the synthesis layer on top. ### Break large inputs into smaller chunks If your dataset is large, don’t feed it all in at once. Split it into smaller pieces that fit comfortably within a context window, process each one separately, and combine the outputs later. ### Break big prompts into smaller tasks A prompt asking for ‘an executive summary, trends, a heatmap, and a narrative for each row’ is actually four or five different tasks. Run them separately. You’ll get better results and it’s easier to spot where something’s gone wrong. ### Consider structured reporting upstream If you’re regularly synthesising data, think about whether you can design the data collection to be more structured in the first place – qualitative and quantitative frameworks that give humans (and AI) cleaner inputs to work from. ## Get back on track Working with data can be complicated and messy, but it’s not as dangerous as slowly drifting away from what’s real. Your strategy or decisions may not break suddenly, but beware the gradual divergence. The antidote isn’t to stop using AI for analysis. It’s to build in the friction that forces you to stay anchored to evidence. Check the outputs. Validate the claims. Don’t let the polished surface of a well-structured AI report substitute for the messier work of knowing your data.

boringmagi.cc

Improving delivery for GOV.‌UK Design System

**Through late 2023 to early 2024, we worked with GOV.‌UK Design System to double their release frequency, significantly improving service delivery speed and team morale.** Facing challenges with traditional two-week sprints, we implemented a new 4-week delivery cycle that fundamentally shifted their approach to rapid iteration and adaptation. ## The problem GOV.‌UK Design System’s mission is to provide reusable components and patterns for government services, ensuring they are usable, accessible, and cost-effective. However, this requires constant iteration and evolution. As technology advances – mobile devices, accessibility standards, and user expectations – the design system needs to adapt at the same pace. Traditional two-week sprints had become tiring and sprint goals often missed, resulting in a slowing down of software releases. This meant the design system often struggled to keep pace with users’ needs for modern, up-to-date components and patterns. ## The approach Recognising this challenge, we adopted a strategic shift towards designing a new delivery model that focuses on delivering value to users, creating space for deep delivery work and learning. This wasn’t just about creating more time for design and development, we also introduced more feedback loops and chances for useful critique. Traditional two-week sprints and Scrum provide good training wheels for teams who are new to agile, but those don’t work for well established or high performing teams. For research and development work (like discovery and alpha), you need a little bit longer to get your head into a domain and have time to play around making scrappy prototypes. For build work, a two-week sprint isn’t really two weeks. With all the ceremonies required for co-ordination and sharing information – which is a lot more labour-intensive in remote-first settings – you lose a couple of days to meetings and workshops. Sprint goals suck too. It’s far too easy to push one along and limp from fortnight to fortnight, never really considering whether you should stop the workstream. It’s better to think about your _appetite_ for doing something, and then to focus on getting valuable iterations out there rather than committing to a whole thing. So we designed a new delivery model to help remove some of these issues. ### A new delivery model The new delivery model is a 4-week cycle split into 3 weeks of delivery or R&D, and 1 week of reflection and planning. You can see how it works in detail on the GOV.‌UK Design System’s team playbook and in a blog post from the team’s delivery manager, Kelly. There are a few principles that make this method work: * Fixed time, variable scope * Think in iterations: vertical not horizontal slices * Each cycle ends with something shippable or showable * R&D cycles end on decisions around scope * Each cycle starts with a brief, but the team has autonomy over delivery This gives space for ideas and conversations to breathe, for spikes and scrappy prototypes to come together, and for teams to make conscious decisions about scope and delivering value to users. The key principle is ‘Fixed time, variable scope’. Teams must end each cycle by either shipping something valuable to users or sharing insights from testing or research. This is what maintains the pace of delivering value consistently. ### Making it work Implementing this new delivery model required close collaboration with other team leads, specifically the delivery manager, Kelly. And it also required a lot of trust from the team to experiment with a new method and reflect on the outcomes. In late 2023, we co-designed the new model with team leads, presenting the principles that would make it work and discussing the activities that would embed key behaviours. After presenting it to the team and getting their buy-in to experiment with a new way of working, the team set off on trying out the new model. ## The impact In their first cycle, the team delivered three out of five briefs – a significant improvement on their completion rate at the time. As Kelly reported, ‘most team members enjoyed working in smaller, focused groups and having autonomy over how they deliver their work.’ A few months later, we analysed how often the team was releasing new software: **they were releasing twice as often in half the time.** Between October 2022 and October 2023, there were five releases. Between October 2023 and March 2024, there were 10 releases. Here’s a graph showing releases per quarter, based on the dates of releases of GOV.‌UK Frontend (their main codebase). (Data collected 28 January 2026.) One year on and the team has maintained momentum. Iterations have increased, they’ve built a steady rhythm of releasing GOV.‌UK Frontend more frequently, and according to a recent review the team is a lot happier due to increased autonomy and reduced frustration with outdated methods. ## Want to try something new? If you’re looking to increase team happiness and effectiveness, look at our services for reviewing your product operating or delivery model and catalysing delivery teams.

boringmagi.cc

Establishing foundations and scaling planning.data.gov.uk

**From March 2024 to March 2026, we worked with the Ministry of Housing, Communities and Local Government to scale its planning data service – restructuring the team, establishing new ways of working, and more than doubling the number of local planning authorities publishing data in the first year alone.** ## The problem The planning data service (planning.data.gov.uk) collects planning and housing data from local planning authorities across England, transforming it into a consistent, open format that anyone can view, download and analyse. Better data leads to better planning decisions, faster development, and ultimately more homes built. But the service had grown organically, and the team structure hadn’t kept pace. Two teams of 27 were responsible for a wide and expanding scope, without clear capability boundaries or shared ways of working. There was no common cadence, no OKR framework, no handbook. The platform had real potential as digital public infrastructure, but the conditions for a growing, distributed team to do its best work weren’t yet in place. ## The approach Joining as Strategic Product Lead, the primary focus was building the conditions for a growing, distributed team to do its best work and deliver on the data platform’s vision at scale. ### Establishing product foundations The first priority was establishing agile cadences from scratch: coordinated two-week sprints, quarterly planning, and fortnightly show-and-tells. An objectives and key results (OKR) framework was introduced that blended leadership and team perspectives, with bi-weekly reviews to track progress and surface blockers early. Key performance indicators (KPIs) and a North Star metric were divised to guide teams’ work and show that the platform is meeting its vision by making it easier to find, use and trust planning and housing data. To make these ways of working durable and shareable, a team handbook was created – documenting how the service is organised, how it works, its product operating model, and the principles that underpin the work. ### Accelerating data provision with a design sprint With those foundations in place, a five-day design sprint was run with the team in London – prototyping a new end-to-end service to make it easier for LPAs to provide data. Testing with people from four LPAs surfaced actionable insights that fed directly into what the team built next. ### Restructuring the team, establishing the operating model Once the ways of working were established, the team structure was redesigned to match the service’s growing scope. Two generalist teams were reorganised into three specialised teams of around 40 people, each aligned to a distinct capability: designing data standards, collecting and managing data, and making data easy to find and use. With the teams settled, an operating model was established to describe how they worked together end to end. Structured around the data lifecycle – from identifying a planning data need through to delivering trusted data on the platform – it gave everyone a shared model of how work flowed across teams, and ensured the platform and its data remained valuable, trustworthy, and aligned to policy and user needs over time. ### Working in the open Working in the open was central to the culture built throughout this period. Public weeknotes, regular blogging, and community drop-ins created a visible trail of evidence that proved its worth when the service came under governance scrutiny as a Government Major Project – involving ministers, senior civil servants, and cross-government partners. ### Prioritising data standards for planning applications and decisions A significant strategic call was prioritising work on data standards for the entire planning permission process – from application submission through to decision. Neither has existed as an open, official, standardised dataset before. The applications standard treats planning applications as structured data for the first time, flowing consistently from applicants through to local authority systems. The decisions standard will create the first open, record-level national dataset of planning outcomes. Together, they lay the foundation for a more transparent planning system – and for new digital services built on top. ### Modelling the economic case To support wider investment decisions, an economic and benefits model was co-developed showing how opening up planning data generates value across the ecosystem: lowering barriers for startups, reducing duplicative data scraping costs for public and private sector organisations, and creating network effects that level the playing field. A Theory of Change projected further transformative benefits, helping to shape business cases for Spending Review. This was borne out by external research – a commissioned report by the UK PropTech Association found the sector has the potential to grow by 20% by 2032 and generate around £72 billion in revenue, with the platform’s open, standardised data as a key enabler. ## The impact The numbers tell a clear story of momentum. The number of LPAs providing at least one dataset grew from 18 to 50 in the first year – and the platform now supports 171 LPAs through the Digital Planning Improvement Fund, with 49 more joining in early 2026. The total number of datasets grew roughly sixfold, and the platform now hosts more than 400 datasets attracting 1.2 million weekly visitors and 20 million monthly hits – four times more than the year before. The design sprint led directly to a better data provision journey, contributing to one of the most significant growth periods in the platform’s history. Local authorities are getting better at publishing higher-quality data more quickly thanks to the support from the service. The platform’s impact has been recognised beyond government. Richard Pope, author of _Platformland_ , highlighted it as an exemplar for UK government digital transformation – a template for creating better digital services and unlocking private sector innovation. Savills and Tract have both publicly highlighted the platform’s transformative potential for the sector. The team handbook and working-in-the-open practices provided the governance evidence needed to navigate Major Project scrutiny. And the prioritisation of the planning applications and decisions data standards represents a landmark moment: for the first time, England will have an open, consistent, record-level dataset covering the entire planning permission process – a foundation for a more transparent, evidence-driven planning system. ## Want to scale your data platform or product teams? If you’re looking to build the conditions for a growing team to do its best work, explore our services for establishing product foundations, designing your OKRs and KPIs, and reviewing your vision or strategy.

boringmagi.cc

Leading AI product innovation on planning.data.gov.uk

**From November 2024, I conceived and coordinated Extract – an AI product that turns decades of trapped planning documents into structured data – then led it through an alpha phase, testing and shaping it with local planning authorities. In June 2025, the Prime Minister launched Extract at London Tech Week and named it one of his AI Exemplars. Early results suggest a task that takes planning officers one to two hours could be completed in under fifteen minutes.** ## The problem England’s planning system runs on data. Modern digital planning tools – used to assess sites, manage applications, and guide development – need high-quality, structured data to function. The problem is that most of that data doesn’t exist in a usable form yet. Across England, over 300 local planning authorities hold decades of planning decisions: conservation area boundaries, Article 4 directions, tree preservation orders, approved planning applications, and more. Most of it is locked in old formats – paper maps, scanned PDFs, microfiche. Without someone manually transcribing it, none of it can be used in a modern digital workflow. The most time-intensive part is creating geospatial data. In one exercise, the Planning Data team found that drawing the boundary shapes for a single conservation area manually took between one and two hours. That’s one document, from one council, for one planning constraint. Multiply it across thousands of documents and hundreds of councils and the scale of the problem becomes hard to ignore. Similar work in the US – digitising restrictive property covenants – was estimated to cost $1.4 million and take 10 years by hand. The government’s housing mission, building 1.5 million homes this Parliament, depends on modern planning infrastructure. But you can’t build modern infrastructure on top of paper. ## The approach ### Spotting the opportunity As a product and design strategist on MHCLG’s Planning Data team, we started conversations with the Incubator for Artificial Intelligence (i.AI) – a team within DSIT that builds AI tools for use across government – about whether AI could help modernise the planning system. We pitched one idea to them early on: using AI to convert historical planning documents into structured data. After those initial conversations, i.AI went away and mapped out over 20 potential directions across the planning domain. When they came back, Extract was at the top of their list – the project they most wanted to work on, having assessed potential impact and feasibility across all of them. The opportunity was unusually clear-cut. The problem was well-documented and wide in scale. The manual effort required was measurable. And there were frontier AI models – vision language models, image segmentation tools – that might, combined in the right way, actually solve it. What mattered wasn’t whether any single technology existed in isolation. It was whether a novel combination of those technologies could do something no existing tool could: extract both text _and_ accurate geospatial data from old maps and documents, at low cost and high speed, in a way that met national planning data standards. ### Setting up the incubation I worked with i.AI to design and set up an incubation: a focused, time-boxed R&D sprint to test whether the core technical hypothesis held. Before any users were involved or any product was built, we needed to know if the underlying technology actually worked well enough to be worth building on. Together, we set specific, measurable success criteria – targets for textual accuracy, date extraction, shape accuracy, and geolocation precision – and agreed guardrails to keep the experiment responsible. The team worked from openly licensed planning documents, built a robust evaluation pipeline, and adopted an evaluation-driven development approach: every model and every line of code tested against measurable targets, iterating systematically rather than chasing a magic solution. You can read the full brief and success criteria on GitHub. Co-ordinating across two government departments with different rhythms, reporting lines, and cultures requires active stewardship. I was the connective tissue: liaising between MHCLG and i.AI, steering when things got stuck, and keeping the strategy coherent across both organisations. In UK government, that kind of cross-departmental co-ordination is harder than it sounds and matters more than it looks. ### What the incubation found The 8-week incubation was extended to 12 weeks when the georeferencing problem proved harder than expected. Eventually it exceeded every success criterion we had set. The team’s biggest invention was a novel method for automatically finding Ground Control Points – identical features on old and modern maps that allow extracted shapes to be accurately located in the real world. There were no prior solutions to draw on. It was genuinely new – real innovation. You can read the full technical account in the MHCLG Digital blog post about the incubation. The efficiency gain was striking. A task taking a planning officer one to two hours could now be completed in under fifteen minutes for approximately 10p. ### The alpha phase With technical feasibility proven, I led the transition into alpha – where a working technical system gets tested with real users to understand what kind of product it should become. I kicked off the alpha, set the guiding principles for the team, and co-ordinated work across MHCLG and i.AI as the product took shape. The team worked to a test-and-learn approach: building prototypes quickly, getting them in front of planning officers and GIS officers at local planning authorities, and using what we learned to change the design. I published weeknotes throughout the alpha phase if you want to follow the work in detail, or check out the design history for screenshots. The core principle was _velocity of learning, not velocity of delivery_. It’s an important distinction for a product at alpha stage. We needed to find out which of our assumptions were wrong before building anything expensive. Research across two rounds of testing with 8 users from 7 local planning authorities showed that the concept landed well. Users grasped what Extract could do almost immediately, and were already thinking about how extracted data could flow into their back-office systems. One participant put it plainly: > “This would definitely speed up the process, but the question is: how transferable is the data to back-office systems?” That kind of response – not just _is it useful_ but _how does it fit into real workflows_ – is exactly what alpha is for. Research also surfaced an unexpected dimension of value. GIS officers at local authorities are typically stretched across many service areas and projects, leaving little time for labour-intensive digitisation work. Users suggested Extract could help by allowing GIS officers to delegate document conversion to colleagues without GIS skills – with the GIS officer reviewing and signing off the output rather than doing the extraction themselves. That’s a different kind of efficiency gain from the one we originally anticipated. Not everything was straightforward. Users found the accuracy of the AI-generated boundary shapes to be a significant factor in their willingness to adopt the output, with many redrawing shapes themselves in the first prototype. That feedback directly shaped the next iteration – adding editing tools, improving the review interface, and exploring ways to produce cleaner polygon outputs. Part of why accuracy matters so much is that planning is a profession, not just a job. The Royal Town Planning Institute’s code of professional conduct requires members to base their advice on reliable evidence and present data clearly and without improper manipulation. That means explainability isn’t just a nice design goal for Extract – it’s a professional obligation for the people using it. A planning officer can’t simply accept an AI-generated boundary and publish it without being able to stand behind it. The ethical and regulatory landscape of the planning domain was a design constraint from the start, not an afterthought. I wrote more about this in ‘So, you want to build AI for professionals’. When the insights indicated we needed more time to learn, we extended the alpha rather than press on prematurely. Trust in the AI’s outputs is a major factor in real-world adoption, and in a domain where professionals are accountable for the data they publish, getting that right matters more than hitting an arbitrary deadline. In February 2026, when my contract was nearing its end, I handed Extract over to another manager to steer the product development sustainably, providing strategic advice and input from the sidelines. The product team are planning for launch in spring 2026. ## The impact In June 2025, the Prime Minister launched Extract at London Tech Week, committing to roll it out across England. It became one of the Prime Minister’s AI Exemplars – a recognition of its potential to deliver at national scale. Extract is still in alpha, transitioning to beta soon, and the clearest measure of impact will come when it reaches councils at scale. But the early signals are strong: * The incubation achieved **100% extraction** of all expected text fields, **94% date accuracy** , and **90% of AI-traced polygon shapes** matching the human-drawn ground truth to a high standard * Testing with local planning authorities in alpha showed Extract reducing document conversion time from around 1 hour to **3–15 minutes per document** – a saving of **75,000 to 95,000 hours** of GIS officer time across 100,000 documents * Even on the first prototype, before editing tools were added, users completed the task in **under 10 minutes** – and some downloaded the data to refine further in their own tools * Georeferencing – automatically locating old map boundaries on a modern map – was identified as the most valued capability, with users describing it as the hardest and most time-consuming part of their current process. As one user put it: _“If you get it wrong with that, there’s a 95% chance of you getting it completely wrong when it comes to the spatial position and accuracy of the feature itself.”_ * Research with **8 users across 7 local planning authorities** found the core concept well understood and the proposition clearly valued * Related data-publishing work at Camden Council showed **60% fewer planning-related enquiries** after publishing clearer data, saving over 21 hours a month – an indication of the downstream value better planning data can unlock The roadmap I put together takes Extract through private and public beta phases, with a live service for local planning authorities planned for 2026. You can read more on i.AI’s project page for Extract. ## Want to do something similar? If you’re looking for someone to identify and incubate high-value AI opportunities, coordinate cross-government teams, or lead an alpha phase for a new public service, get in touch.

boringmagi.cc

Recap of the AI Leaders panel, January 2026

<p>Here’s what we said at Ministry of Justice’s <a href="https://events.teams.microsoft.com/event/de1b76ef-757f-4fcd-9ef6-c98d4e480467@c6874728-71e6-41fe-a9e1-2e8c36776ad8">AI Leaders panel</a>, the opening event of their <a href="https://www.eventbrite.com/e/ai-digital-professions-conference-tickets-1970192031402?viewDetails=true">AI Digital Professions conference</a> in January 2026.</p> <p>Thank you to Nikola Goger for the invitation to speak, and thanks too to the other panellists: Tim Paul (Head of User Centred Design at i.AI), Paul Haigh (Chief Technology Officer at Ministry of Justice) and Jan Murdoch (Head of Horizon Scanning at Department for Environment, Food &amp; Rural Affairs).</p> <p>This is a tidied-up version of what was said – there were many more instances of ‘um’, ‘like’ and far too many instances of ‘I think’ in the original transcript!</p> <h3 id="question-how-are-you-seeing-ai-change-each-of-your-professions-or-areas-and-how-do-you-think-this-will-impact-the-future-of-your-profession-or-area">Question: How are you seeing AI change each of your professions or areas? And how do you think this will impact the future of your profession or area?</h3> <p>As the product person, I’ve taken a broader team view rather than focusing solely on product-specific matters. However, I believe the key point Mark highlighted earlier is that writing code isn’t necessarily a bottleneck anymore. This raises questions about how it impacts the rest of your process and overall workflow.</p> <p>Essentially, it shifts that bottleneck earlier in the value stream. Consequently, you must prioritise establishing solid foundations for product architecture and design. For example, if agents can easily complete tickets quickly, you must ensure these foundations are in place to deliver successful products and services.</p> <p>There’s a slight advantage to it: delivery could potentially become faster. However, there’s also a risk of delivering the wrong thing. Therefore – if we consider service stages – the alpha stage really should be seen as a way to test various solutions and concepts, really try out different ideas.</p> <p>You have the ability to explore more divergent thinking, which is a good thing. Therefore, you want to optimise your approach to cycle around the ‘problem understanding, solution exploration’ loop more quickly and more often. This means using design methods that encourage this divergence and generate a variety of ideas, and then methods of testing that give you clear signals on where to converge.</p> <p>I think the impact lies in the fact that if we’ve been discussing agile practices for the past decade, we’re transitioning towards a phase where Lean practices and Lean Startup methods will become significantly more integral to our work. Mastering writing hypotheses and measuring outcomes will be crucial. Furthermore embracing experimentation and remaining open to being proven wrong will be incredibly beneficial.</p> <p>Now that computers can do the some of the computer tasks, we should look at shifting our focus to human-centric activities. We should look more at using co-design, which offers efficiencies in prototyping and testing. Right now you test a product with five people on one day, make some iterations, then return a few days later to test again. Instead, we could gather groups of people in a room for a few days, allowing them to iterate the product with us and provide valuable feedback. This approach empowers users and generates stronger evidence. (And as many people know, access to users can be a stultifying constraint.)</p> <p>You know, in the product profession we consistently emphasise falling in love with the problem as an important thing to do. Interestingly, this concept resonates in the private sector too. For example, Y Combinator’s Paul Graham frequently discusses that those startups who secure funding and Y Combinator’s are the ones who are intensely focused on users and obsessed with their problems. This aligns closely with our longstanding focus on prioritising user needs in government digital services. Ultimately, truly understanding the problem and having skilled designers who grasp constraints and engineers who design effective architecture will be crucial in delivering high-quality experiences, especially now more than ever.</p> <p>We need to become much more strategic and excel at demonstrating outcomes rather than simply pumping out a digital version of a form.</p> <h3 id="question-lets-talk-limits-where-do-you-see-the-limits-of-ai-in-public-service-delivery-specifically-so-what-are-the-things-that-should-never-be-automated-and-what-kind-of-unintended-consequences-have-you-seen-or-you-think-are-there">Question: Let’s talk limits. Where do you see the limits of AI in public service delivery specifically? So what are the things that should never be automated and what kind of unintended consequences have you seen or you think are there?</h3> <p>Tim’s point about judgement is really important. It’s something we think about working on with Extract, a product that uses AI to extract geospatial data from documents. I did go down a bit of a rabbit hole looking at automated decision-making and how does AI have an impact on professionals?</p> <p>When you consider professionals like doctors and lawyers, you’re looking at the professionalisation of judgement and decision-making. The way a profession is structured, including its educational pathway, specific behaviours and codes of conduct, this all revolves around explainability and audit-ability. This allows for an in-depth examination and inspection of how decisions were made, meaning you can adjudicate what happened. However, stuffing AI into things introduces a ‘black box’ element, you lose that transparency. And while there are people actively working on explainability and shining a light into the black box, I believe it’s a crucial aspect in considering what not to automate.</p> <p>Returning to our digital professions, I wanted to discuss product management. I believe a potential danger arises when AI is used for tasks like writing tickets, vision statements or mission briefs. As I joked with Tim a few years ago, product management essentially boils down to prompt engineering! This is partly true because providing sufficient context and direction is crucial for achieving successful outcomes. I still believe this is the core responsibility of the role.</p> <p>Therefore, ensuring everyone has a comprehensive understanding of users, their context and the problem space they’re working within, is essential. This, I think, is the true craft of product management. Much of this occurs in tickets, in discussions, and vision statements, among other things. Therefore, I don’t think outsourcing that craft is a good idea.</p> <p>Some times I’m okay with it, like writing boring tickets. For example, recently I had to create tickets for implementing an analytics service on some prototype. I went to look at the documentation and realised I didn’t need to do it myself anymore. I could simply point an agent at the documentation and ask them to write a decent ticket. All I had to do was outline the stages I wanted to measure and the metrics I wanted to capture. But the agent was able to write the ticket for me in a very short space of time.</p> <p>The danger with using AI in some product jobs is that it’s all about managing context. You need to be careful about what you include in people’s context, and that’s really the craft of product management. So be present, don’t automate that.</p> <h3 id="question-what-skill-sets-do-digital-professionals-need-to-work-effectively-with-ai-in-your-profession-and-what-would-you-advise-people-to-learn">Question: What skill sets do digital professionals need to work effectively with AI in your profession? And what would you advise people to learn?</h3> <p>I mentioned it earlier but I think improving your interpersonal skills – the human stuff – is really beneficial. This includes developing conversation skills, getting better at writing and the ability to bring groups together. Get used to considering and discussing feedback together, having differing points of view.</p> <p>I believe we should create more space and time for apprenticeships. Earlier, Paul mentioned hiring <em>more</em> juniors rather than fewer, which is great. It reminds me of Yanagi Soetsu, an arts and crafts historian from the early 20th century. During the period of mechanisation in Korea and Japan, he discussed the importance of preserving craft. He talked about ‘village kilns’ and emphasised the importance of skilled artisans working with apprentices to teach methods and pass on knowledge. I believe this will be incredibly crucial in the coming years, otherwise what’s the social contract we’re signing up to?</p> <p>At a day-to-day level, I believe simply downloading some open large language models and using them through software like LM Studio will be incredibly beneficial. Start by understanding the differences between various models, observing their responses to prompts and how tweaking them produces different outputs. Learn to control inference, such as extending the context window, and delve into data and statistics. I genuinely think this knowledge will be very handy.</p> <p>Jeremy Keith once said AI is simply applied statistics, which is both funny and true. Therefore, if you learn the feel of the material, you can get better at using it (or not).</p>

boringmagi.cc

Metrics, measures and indicators: a few things to bear in mind

Metrics, measures and indicators help you track and evaluate outcomes. They can tell us if we’re moving in the right direction, if things aren’t going well, or if we’ve achieved the outcome we set out to achieve. If you’ve reported on key performance indicators (KPIs), checked progress against objectives and key results (OKRs) or looked at user analytics, you’ll have some experience with metrics, measures and indicators. These words are often used interchangably and, in general, the difference isn’t important. Not for this post anyway. We can talk about the difference between metrics, measures and indicators later. In this post we’ll cover some guiding principles for designing and using metrics, measures and indicators. A few things to bear in mind. ## Guiding principles 1. Value outcomes over outputs 2. Measures, not targets 3. Balance the what (quantitative) and the why (qualitative) 4. Measure the entire product or service 5. Keep it light and actionable 6. Revisit or refine as things change ### Value outcomes over outputs We acknowledge that outputs are on the path to achieving outcomes. You can’t cater for a memorable birthday party without making some sandwiches. But delivering outcomes is the real reason why we’re here. So we don’t measure whether we’ve delivered a product or feature, we measure the impact it’s having. ### Measures, not targets Follow Goodhart’s Law: ‘When a measure becomes a target, it ceases to be a good measure.’ There are numerous factors that contribute to a number or reading going up or down. Metrics, measures and indicators are a starting point for a conversation, so we can ask why and do something about it (or not). The measures are in service of learning: tools, not goals. ### Balance the what (quantitative) and the why (qualitative) Grown-ups love numbers. But it’s very easy to ignore what users think and feel when you only track quantitative measures. Numbers tell us what’s happening, but feedback can tell us why. There’s no point doing something faster if it makes the experience worse for users, for example – we have to balance quantity and quality. ### Measure the entire product or service If we can see where people start, how they move through and where they end, we can identify where to focus our efforts for improvements. The same is true for people who come back too, we want to see whether we’ve made things better than last time they were here. If you’re only measuring one part, you only know how one part is performing. Get holistic readings. ### Keep them light and actionable It’s easy to go overboard and start tracking everything, but too much information can be a bad thing. If we track too many metrics, we run the risk of analysis paralysis. Similarly, one measure is too few: it’s not enough to understand an entire system. Four to eight key metrics or indicators per team is enough and should inspire action. ### Revisit or refine as things change Our priorities will change over time, meaning we will need to change our indicators, measures and metrics too. It’s no use tracking and reporting on datapoints that don’t relate to outcomes. Measure what matters. We should aim not to change them too frequently – that causes whiplash. But it’s all right to change them when you change direction or focus. ## Are we on the way? Or did we get there? Those principles are handy for working out what to measure, but there’s two types of indicator you need to know about: leading and lagging. Leading indicators tell us whether we’re making progress towards an outcome. _Are we on the way?_ For example, if we want to make it easy to find datasets, are people searching for data? Is the number of people searching for data going up? Lagging indicators tell us whether we’ve achieved the outcome. _Did we get there?_ In that same example, making it easy to find datasets, what’s the user satisfaction score? Are they requesting new datasets?

boringmagi.cc

Using quarters as a checkpoint

Breaking your strategy down into smaller, more manageable chunks can help you make more progress sooner. Some things take a long while to achieve, but smaller goals help us celebrate the wins along the way. Many organisations use a quarter – a block of 3 months – to do this. And it can be helpful to look back before you look forward, to celebrate the progress you’ve made and work out what to do next. Every 3 months, we encourage product teams to take the opportunity to step back from the day-to-day and consider the objectives they’re working towards. The quarterly checkpoint is a time to refocus efforts and double-down, change direction or move on to the next objective. There are 2 stages to using the quarterly checkpoint well: 1. Check on your progress 2. Plan how to achieve your new objectives Here are two workshops you can run at each stage, but you can combine them into one workshop if you like. Whatever works. ## Check on your progress First, check on the progress your team has made on your objectives and key results (OKRs). You can do this in a team workshop lasting 30 to 60 minutes. ### 1. List out the OKRs you’ve been working on (10 to 20 mins) Run through the OKRs you’ve been working on. Talk about the progress you made on each key result and celebrate the successes – big or small! ### 2. Think about what’s left to do (20 to 40 mins) For any OKRs you haven’t completed – where progress on key results isn’t 100% – discuss as a team which initiatives you have left to do to fully achieve the objective. For example, you may need to collect some data, run a test, build a thing or achieve an outcome. Consider whether you should change your approach, for example, by doing something smaller or using different methods, based on what you’ve learned over the last quarter. It’s OK to stick to the original plan if it’s still the best approach. Write down what initiatives your team has agreed to do. ## Plan how to achieve your new objectives Next, you’ll need to form a loose plan for how to achieve your new objectives. You can treat unfinished objectives from the previous quarter as a new objective. Run another workshop lasting 30 to 45 minutes for each objective. Everyone on the team will need to input on the plan using the outline below. Write it in a doc, a slide deck or on a whiteboard – whatever works. You will probably want to present these plans to the senior management team or service owner at the start of the new quarter. If it’s easier than starting with a blank page, team leads can fill in the outline and get feedback from the rest of the team. As long as everyone gets a chance to input, it doesn’t matter. It’s OK if you take less than 30 minutes, especially if you already have a plan. ### 1. Write down and describe the objective An objective is a bold and qualitative goal that the organisation wants to achieve. It’s best that they’re ambitious, not super easy to achieve or audacious in nature; they are not sales targets. Write down the problem you’re solving and who it’s a problem for. Discuss how you’ll know when you’re done. What are the success criteria? ### 2. Think about risks and unknowns What might be a challenge? What are the riskiest assumptions or big unknowns to highlight? Do you need to try new techniques? These might form the first initiatives in your plan. You can frame your assumptions using the hypothesis statement: **Because** [of something we think or know] **We believe that** [an initiative or task] **Will achieve** [desired outcome] **Measured by** [metric] Note down dependencies on other teams, for example, where you may need another team to do something for you. ### 3. Detail all the initiatives Write a sentence for all the initiatives – tasks and activities – you’ll need to do to achieve the objective. Consider research and discovery activities, which can help you gather information to turn unknowns into knowns. Consider alphas, things to prototype, spikes, and experiments that can help you de-risk or validate assumptions. Make sure to remember the development and delivery work too – that’s how we release value to users! ### 4. What will you measure? Review your success criteria. Define the metrics that will tell you when you’ve finished or achieved the objective. These should tell you when you’re done and will become your key results. Remember, metrics should be: * tangible and quantitative * specific and measurable * achievable and realistic ### 5. Prioritise radically What would you do differently if you only had half the time? How will you start small and build up? What’s the least amount of work you can do to learn the most? Use these thoughts to consider any changes to your initiatives. Go back and edit the initiatives if you need to. ## Don’t worry about adapting your plans A core tenet of agile is responding to change over following a plan, so don’t be afraid to change your plans based on new information. The quarterly checkpoint isn’t the only time you can look back to look forward – that’s why retrospectives are useful. You can use the activities above at any point. The best product teams build these behaviours into their regular practice. If you’d like help running these workshops or have any questions, get in touch and we’ll set up a chat.

boringmagi.cc

Going faster with Lean principles

Software teams are often asked to go faster. There are many factors that influence the speed at which teams can discover, design and deliver solutions, and those factors aren’t always in a team’s control. But Lean principles offer teams a way to analyse and adapt their operating model – their ways of working. ## What is Lean? Lean is a method of manufacturing that emerged from Toyota’s Production System in the 1950s and 1960s. It’s a system that incorporates methods of production and leadership together. The early Agile community used Lean principles to inspire methods for making digital products and services. These principles have had influence beyond the production environment and have been adapted for business and strategy functions too. ## Books on Lean Four books on Lean principles have influenced the way I work. **1._Lean Software Development: An Agile Toolkit_ by Mary and Tom Poppendieck** The earliest of the four books. It really set the standard. **2._The Lean Startup_ by Eric Ries** This started a big movement for applying Lean principles to your startup, including testing out new business models or growth opportunities. **3._Lean UX_ by Jeff Gothelf and Josh Seiden** One of my favourites. This one really brought strategic goals and user experience closer together. It also shifted teams from writing problem statements to writing hypotheses. **4._The Lean Product Playbook_ by Dan Olsen** This is relatively similar to _The Lean Startup_ but is more of a playbook, showing the practice that goes with the theory. The highlight is its emphasis on MVP tests: experiments you can run to learn something without building anything. ## Lean principles All these books have some principles in their pages, all based on the original Lean principles from Toyota. They’re all pretty similar. Combining their approaches helps us apply Lean principles to business model development, strategy, user-centred design and software delivery. > A note on principles: Principles are not rules. Principles guide your thinking and doing. Rules say what’s right and wrong. ### 1. Eliminate waste Reduce anything which does not help deliver value to the user. So: partially done work; scope creep; re-learning; task-switching; waiting; hand-offs; defects; management activities. Outcomes, not outputs. ### 2. Amplify learning Build, measure, learn. Create feedback loops. Build scrappy prototypes, run spikes. Write tests first. Think in iterations. ### 3. Decide as late as possible Call out the assumptions or uncertainties, try out different options, and make decisions based on facts or evidence. ### 4. Deliver as fast as possible Shorter cycles improve learning and communication, and helps us meet users’ needs as soon as possible. Reduce work in progress, get one thing done, and iterate. ### 5. Empower the team Figure it out together. Managers provide goals, encourage progress, spot issues and remove impediments. Designers, developers and data engineers suggest how to achieve a goal and feed in to continuous improvement. ### 6. Build integrity in Agility needs quality. Automated tests and proven design patterns allow you to focus on smaller parts of the system. A regular flow of insights to act on aids agility. ### 7. Optimise the whole Focus on the entire value stream, not just individual tasks. Align strategy with development. Consider the entire user experience in the design process. ## Three simpler principles If those seem like too many to get started with, I want to introduce three simpler principles that can help you go faster. I came across these in a book about running, which doesn’t seem like the place you’d find inspiration about product management! Think easy, light and smooth. It’s from a man called Micah True who lived in the Mexican desert and went running with the local Native Americans. They called him Caballo Blanco – ‘White Horse’ – because of his speed. > “You start with easy, because if that’s all you get, that’s not so bad. Then work on light. Make it effortless, like you don’t give a shit how high the hill is or how far you’ve got to go. When you’ve practised that so long that you forget you’re practicing, you work on making it smooooooth. You won’t have to worry about the last one – you get those three, and you’ll be fast.” You can do this every cycle. Find one thing to make easier, one thing to make lighter, and one thing to make smoother. Fast will happen naturally.

boringmagi.cc

Our positions on generative AI

Like many trends in technology before it, we’re keeping an eye on artificial intelligence (AI). AI is more of a concept, but generative AI as a general purpose technology has come to the fore due to recent developments in cloud-based computation and machine learning. Plus, technology is more widespread and available to more people, so more people are talking about generative AI – compared to something _even more_ ubiquitous like HTML. Given the hype, it feels worthwhile stating our positions on generative AI – or as we like to call it, ‘applied statistics’. We’re open to working on and with it, but there’s a few ideas we’ll bring to the table. ## The positions 1. Utility trumps hyperbole 2. Augmented not artificial intelligence 3. Local and open first 4. There will be consequences 5. Outcomes over outputs ### Utility trumps hyperbole The fundamental principle to Boring Magic’s work is that people want technologies to work. People prefer things to be functional first; the specific technologies only matter when they reduce or undermine the quality of the utility. There are outsized, unfounded claims being made about the utility of AI. It is not ‘more profound than fire’. The macroeconomic implications of AI are often overstated, but it’ll still likely have an impact on productivity. We think it’s sensible to look at how generative AI can be useful or make things less tedious, so we’re exploring the possibilities: from making analysis more accessible through to automating repeatable tasks. We won’t sell you a bunch of hype, just deliver stuff that works. ### Augmented not artificial intelligence Technologies have an impact on the availability of jobs. The introduction of the digital spreadsheet meant that chartered accountants could easily punch the numbers, leading to accounting clerks becoming surplus to requirements. Jevon’s paradox teaches us that AI will lead to more work, not less. Over time accountants needed fewer clerks, but increases in financial activity have lead to a greater need for auditors. So we will still need people in jobs to do thinking, reasoning, assessing and other things people are good at. Rather than replacing people with machines to reduce costs, technology should be used to empower human workers. We should augment the intelligence of our people, not replace it. That means using things like large language models (LLMs) to reduce the inertia of the blank page problem, helping with brainstorming, rather than asking an LLM to write something for you. Extensive not intensive technology. ### Local and open first Right now, we’re in a hype cycle, with lots of enthusiasm, funding and support for generative AI. The boom of a hype cycle is always followed by a bust, and AI winters have been common for decades. If you add AI to your product or service and rely on a cloud-based supplier for that capability, you could find the supplier goes into administration – or worse, enshittification, when fees go up and the quality of service plunges. And free services are monetised eventually. But there are lots of openly-available generative text and vision models you can run on your own computer – your ‘local machine’ – breaking the reliance on external suppliers. When exploring how to apply generative AI to a client’s problem, we’ll always use an open model and run it locally first. It’s cheaper than using a third party, and it’s more sustainable too. It also mitigates some risks around privacy and security by keeping all data processing local, not running on a machine in a data centre. That means we can get started sooner and do a data protection impact assessment later, when necessary. We can use the big players like OpenAI and Anthropic if we need to, but let’s go local and open first. ### There will be consequences People like to think of technology as a box that does a specific thing, but technology impacts and is impacted by everything around it. Technology exists within an ecology. It’s an inescapable fact, so we should try to sense the likely and unlikely consequences of implementing generative AI – on people, animals, the environment, organisations, policy, society and economies. That sounds like a big project, but there are plenty of tools out there to make it easier. We’ve used tools like consequence scanning, effects mapping, financial forecasting, Four Futures and other extrapolation methods to explore risks and harms in the past. As responsible people, it’s our duty to bring unforeseen consequences more into view, so that we can think about how to mitigate the risks or stop. ### Outcomes over outputs It feels like everyone’s doing something with generative AI at the moment, and, if you’re not, it can lead to feeling left out. But this doesn’t mean you have to do something: FOMO is not a strategy. We’ll take a look at where generative AI might be useful, but we’ll also recommend other technologies if those are cheaper, faster or more sustainable. That might mean implementing search and filtering instead of a chatbot, especially if it’s an interface that more people are used to. It’s more important to get the job done and achieve outcomes, instead of doing the latest thing because it’s cool. ## Let’s be pragmatic Ultimately our approach to generative AI is like any other technology: we’re grounded in practicality, mindful of being responsible and ethical, and will pursue meaningful outcomes. It’s the best way to harness its potential effectively. Beware the AI snake oil.

boringmagi.cc

Tips on doing show & tell well

## What is a show & tell? A show & tell is a regular get-together where people working on a product or service celebrate their work, talk about what they’ve learned, and get feedback from their peers. It’s also a chance to * bring together team members, management and leadership to bond, share success, and collaborate * let colleagues know what you’re working on, keep aligned, and create opportunities to connect and work together * tell stakeholders (including users, partner organisations and leadership) what you’ve been doing and take their questions as feedback (a form of governance). A show & tell may be internal, limited to other people in the same team or organisation, or open to anyone to join. Most teams start with an internal show & tell and make these open later. A show & tell might also be called a team review. ## How to run a great show & tell 1. **Don’t make it up on the spot** Spend time as a team working out what you want to say and who is going to share stories with the audience (1 or 2 people works best). 30 to 60 minutes of prep will pay off. 2. **Set the scene** Always introduce your project or epic. Who’s on the team? What are you working on? What problem are you solving? Who are your users? Why are you doing it? You don’t need to tell the full history, a 30-second overview is enough. 3. **Show the thing!** Scrappy diagrams, Mural boards, Post-it notes, screenshots, scribbles, photos, and clicking through prototypes bring things to life. Text and code is OK, but always aim to demonstrate something working – don’t just talk through a doc or some function. 4. **Talk about what you’ve learned** Share which assumptions turned out to be incorrect, or what facts surprised you. Show clips from user research and usability testing. Highlight important analytics data or performance measures. Share both findings and insights. Be clear on the methodology and any confidence intervals, levels of confidence, risky assumptions, etc. 5. **Be clear** Don’t hide behind jargon. Make bold statements. Say what you actually think! This helps everyone concentrate on the main point, and it generates discussion. 1. **Always share unfinished thinking** Forget about the polish and perfection. A show & tell is the perfect place to collect feedback, ideas and thoughts. It’s a complicated space. We’re all trying to figure it out! 2. **Rehearse** Take 10–15 minutes to rehearse your section with your team to work out whether you need to cut anything. If you’re struggling to edit, use a format like What? So what? Now what? to keep things concise. If you take up more time than you’ve been given, it’ll eat into other people’s section meaning they have to rush (or not share at all) which isn’t fair. 3. **Leave time for questions** The best show & tells have audience participation. Wherever possible, leave time for questions – either after each team or at the end. Encourage people to ask questions in the chat, on Slack, in docs, etc. If you do nothing else, follow tip number 3. You can read more tips on good show & tells from Mark Dalgarno, Emily Webber and Alan Wright. ## How to be a great show & tell audience member 1. **Be present and listen** There’s nothing worse than preparing for a show & tell only to realise that no one’s paying attention. Close Slack, close Teams, stop looking at email, and give your full attention to your team-mates. 2. **Smile, use emojis, and celebrate!** Bring the good vibes and lift each other up whenever there’s something worth celebrating. ## It’s ok to be halfway done The main thing to remember is that show & tell is not just about sharing progress and successes. It’s a time to talk about what’s hard and what didn’t work too. It’s ok to be halfway done. It’s ok to go back to the drawing board. Each sprint, try to answer these questions in your show & tell: * What did we learn or what changed our mind? * What can we show? How can we help people see behind the scenes? * What haven’t we figured out? What do we want feedback on?

boringmagi.cc

You don’t have to do fortnightly sprints

In early 2024, we helped GOV.‌UK Design System design and implement a new model for agile delivery. It was a break away from traditional Scrum and two-week sprints towards an emphasis on iteration and reflection. ## Why change things? Traditional two-week sprints and Scrum provide good training wheels for teams who are new to agile, but those don’t work for well established or high performing teams. For research and development work (like discovery and alpha), you need a little bit longer to get your head into a domain and have time to play around making scrappy prototypes. For build work, a two-week sprint isn’t really two weeks. With all the ceremonies required for co-ordination and sharing information – which is a lot more labour-intensive in remote-first settings – you lose a couple of days with two-week sprints. Sprint goals suck too. It’s far too easy to push it along and limp from fortnight to fortnight, never really considering whether you should stop the workstream. It’s better to think about your appetite for doing something, and then to focus on getting valuable iterations out there rather than committing to a whole thing. ## How it works You can see how it works in detail on the GOV.‌UK Design System’s team playbook and in a blog post from the team’s delivery manager, Kelly. There’s also a graphic that brings the four-week cycle to life. There are a few principles that make this method work: * Fixed time, variable scope * Think in iterations: vertical not horizontal slices * Each cycle ends with something shippable or showable * R&D cycles end on decisions around scope * Each cycle starts with a brief, but the team has autonomy over delivery This gives space for ideas and conversations to breathe, for spikes and scrappy prototypes to come together, and for teams to make conscious decisions about scope and delivering value to users. ## How did it work out? In their first cycle, the team delivered three out of five briefs – which was higher than their completion rate at the time. As Kelly reported, ‘most team members enjoyed working in smaller, focused groups and having autonomy over how they deliver their work.’ A few months later, we analysed how often the team was releasing new software: **they were releasing twice as often in half the time.** Between October 2022 and October 2023, there were five releases. Between October 2023 and March 2024, there were 10 releases. One year on and the team has maintained momentum. Iterations have increased, they’ve built a steady rhythm of releasing GOV.‌UK Frontend more frequently, and according to a recent review the team is a lot happier working that way. ## Want to try something new? If you’re looking to increase team happiness and effectiveness, drop us a line and we can chat about transforming your team’s delivery model too.

boringmagi.cc

Using quarters as a checkpoint

Breaking your strategy down into smaller, more manageable chunks can help you make more progress sooner. Some things take a long while to achieve, but smaller goals help us celebrate the wins along the way. Many organisations use a quarter – a block of 3 months – to do this. And it can be helpful to look back before you look forward, to celebrate the progress you’ve made and work out what to do next. Every 3 months, we encourage product teams to take the opportunity to step back from the day-to-day and consider the objectives they’re working towards. The quarterly checkpoint is a time to refocus efforts and double-down, change direction or move on to the next objective. There are 2 stages to using the quarterly checkpoint well: 1. Check on your progress 2. Plan how to achieve your new objectives Here are two workshops you can run at each stage, but you can combine them into one workshop if you like. Whatever works. ## Check on your progress First, check on the progress your team has made on your objectives and key results (OKRs). You can do this in a team workshop lasting 30 to 60 minutes. ### 1. List out the OKRs you’ve been working on (10 to 20 mins) Run through the OKRs you’ve been working on. Talk about the progress you made on each key result and celebrate the successes – big or small! ### 2. Think about what’s left to do (20 to 40 mins) For any OKRs you haven’t completed – where progress on key results isn’t 100% – discuss as a team which initiatives you have left to do to fully achieve the objective. For example, you may need to collect some data, run a test, build a thing or achieve an outcome. Consider whether you should change your approach, for example, by doing something smaller or using different methods, based on what you’ve learned over the last quarter. It’s OK to stick to the original plan if it’s still the best approach. Write down what initiatives your team has agreed to do. ## Plan how to achieve your new objectives Next, you’ll need to form a loose plan for how to achieve your new objectives. You can treat unfinished objectives from the previous quarter as a new objective. Run another workshop lasting 30 to 45 minutes for each objective. Everyone on the team will need to input on the plan using the outline below. Write it in a doc, a slide deck or on a whiteboard – whatever works. You will probably want to present these plans to the senior management team or service owner at the start of the new quarter. If it’s easier than starting with a blank page, team leads can fill in the outline and get feedback from the rest of the team. As long as everyone gets a chance to input, it doesn’t matter. It’s OK if you take less than 30 minutes, especially if you already have a plan. ### 1. Write down and describe the objective An objective is a bold and qualitative goal that the organisation wants to achieve. It’s best that they’re ambitious, not super easy to achieve or audacious in nature; they are not sales targets. Write down the problem you’re solving and who it’s a problem for. Discuss how you’ll know when you’re done. What are the success criteria? ### 2. Think about risks and unknowns What might be a challenge? What are the riskiest assumptions or big unknowns to highlight? Do you need to try new techniques? These might form the first initiatives in your plan. You can frame your assumptions using the hypothesis statement: **Because** [of something we think or know] **We believe that** [an initiative or task] **Will achieve** [desired outcome] **Measured by** [metric] Note down dependencies on other teams, for example, where you may need another team to do something for you. ### 3. Detail all the initiatives Write a sentence for all the initiatives – tasks and activities – you’ll need to do to achieve the objective. Consider research and discovery activities, which can help you gather information to turn unknowns into knowns. Consider alphas, things to prototype, spikes, and experiments that can help you de-risk or validate assumptions. Make sure to remember the development and delivery work too – that’s how we release value to users! ### 4. What will you measure? Review your success criteria. Define the metrics that will tell you when you’ve finished or achieved the objective. These should tell you when you’re done and will become your key results. Remember, metrics should be: * tangible and quantitative * specific and measurable * achievable and realistic ### 5. Prioritise radically What would you do differently if you only had half the time? How will you start small and build up? What’s the least amount of work you can do to learn the most? Use these thoughts to consider any changes to your initiatives. Go back and edit the initiatives if you need to. ## Don’t worry about adapting your plans A core tenet of agile is responding to change over following a plan, so don’t be afraid to change your plans based on new information. The quarterly checkpoint isn’t the only time you can look back to look forward – that’s why retrospectives are useful. You can use the activities above at any point. The best product teams build these behaviours into their regular practice. If you’d like help running these workshops or have any questions, get in touch and we’ll set up a chat.

boringmagi.cc

Tips on doing show & tell well

## What is a show & tell? A show & tell is a regular get-together where people working on a product or service celebrate their work, talk about what they’ve learned, and get feedback from their peers. It’s also a chance to * bring together team members, management and leadership to bond, share success, and collaborate * let colleagues know what you’re working on, keep aligned, and create opportunities to connect and work together * tell stakeholders (including users, partner organisations and leadership) what you’ve been doing and take their questions as feedback (a form of governance). A show & tell may be internal, limited to other people in the same team or organisation, or open to anyone to join. Most teams start with an internal show & tell and make these open later. A show & tell might also be called a team review. ## How to run a great show & tell 1. **Don’t make it up on the spot** Spend time as a team working out what you want to say and who is going to share stories with the audience (1 or 2 people works best). 30 to 60 minutes of prep will pay off. 2. **Set the scene** Always introduce your project or epic. Who’s on the team? What are you working on? What problem are you solving? Who are your users? Why are you doing it? You don’t need to tell the full history, a 30-second overview is enough. 3. **Show the thing!** Scrappy diagrams, Mural boards, Post-it notes, screenshots, scribbles, photos, and clicking through prototypes bring things to life. Text and code is OK, but always aim to demonstrate something working – don’t just talk through a doc or some function. 4. **Talk about what you’ve learned** Share which assumptions turned out to be incorrect, or what facts surprised you. Show clips from user research and usability testing. Highlight important analytics data or performance measures. Share both findings and insights. Be clear on the methodology and any confidence intervals, levels of confidence, risky assumptions, etc. 5. **Be clear** Don’t hide behind jargon. Make bold statements. Say what you actually think! This helps everyone concentrate on the main point, and it generates discussion. 1. **Always share unfinished thinking** Forget about the polish and perfection. A show & tell is the perfect place to collect feedback, ideas and thoughts. It’s a complicated space. We’re all trying to figure it out! 2. **Rehearse** Take 10–15 minutes to rehearse your section with your team to work out whether you need to cut anything. If you’re struggling to edit, use a format like What? So what? Now what? to keep things concise. If you take up more time than you’ve been given, it’ll eat into other people’s section meaning they have to rush (or not share at all) which isn’t fair. 3. **Leave time for questions** The best show & tells have audience participation. Wherever possible, leave time for questions – either after each team or at the end. Encourage people to ask questions in the chat, on Slack, in docs, etc. If you do nothing else, follow tip number 3. You can read more tips on good show & tells from Mark Dalgarno, Emily Webber and Alan Wright. ## How to be a great show & tell audience member 1. **Be present and listen** There’s nothing worse than preparing for a show & tell only to realise that no one’s paying attention. Close Slack, close Teams, stop looking at email, and give your full attention to your team-mates. 2. **Smile, use emojis, and celebrate!** Bring the good vibes and lift each other up whenever there’s something worth celebrating. ## It’s ok to be halfway done The main thing to remember is that show & tell is not just about sharing progress and successes. It’s a time to talk about what’s hard and what didn’t work too. It’s ok to be halfway done. It’s ok to go back to the drawing board. Each sprint, try to answer these questions in your show & tell: * What did we learn or what changed our mind? * What can we show? How can we help people see behind the scenes? * What haven’t we figured out? What do we want feedback on?

boringmagi.cc

Going faster with Lean principles

Software teams are often asked to go faster. There are many factors that influence the speed at which teams can discover, design and deliver solutions, and those factors aren’t always in a team’s control. But Lean principles offer teams a way to analyse and adapt their operating model – their ways of working. ## What is Lean? Lean is a method of manufacturing that emerged from Toyota’s Production System in the 1950s and 1960s. It’s a system that incorporates methods of production and leadership together. The early Agile community used Lean principles to inspire methods for making digital products and services. These principles have had influence beyond the production environment and have been adapted for business and strategy functions too. ## Books on Lean Four books on Lean principles have influenced the way I work. **1._Lean Software Development: An Agile Toolkit_ by Mary and Tom Poppendieck** The earliest of the four books. It really set the standard. **2._The Lean Startup_ by Eric Ries** This started a big movement for applying Lean principles to your startup, including testing out new business models or growth opportunities. **3._Lean UX_ by Jeff Gothelf and Josh Seiden** One of my favourites. This one really brought strategic goals and user experience closer together. It also shifted teams from writing problem statements to writing hypotheses. **4._The Lean Product Playbook_ by Dan Olsen** This is relatively similar to _The Lean Startup_ but is more of a playbook, showing the practice that goes with the theory. The highlight is its emphasis on MVP tests: experiments you can run to learn something without building anything. ## Lean principles All these books have some principles in their pages, all based on the original Lean principles from Toyota. They’re all pretty similar. Combining their approaches helps us apply Lean principles to business model development, strategy, user-centred design and software delivery. > A note on principles: Principles are not rules. Principles guide your thinking and doing. Rules say what’s right and wrong. ### 1. Eliminate waste Reduce anything which does not help deliver value to the user. So: partially done work; scope creep; re-learning; task-switching; waiting; hand-offs; defects; management activities. Outcomes, not outputs. ### 2. Amplify learning Build, measure, learn. Create feedback loops. Build scrappy prototypes, run spikes. Write tests first. Think in iterations. ### 3. Decide as late as possible Call out the assumptions or uncertainties, try out different options, and make decisions based on facts or evidence. ### 4. Deliver as fast as possible Shorter cycles improve learning and communication, and helps us meet users’ needs as soon as possible. Reduce work in progress, get one thing done, and iterate. ### 5. Empower the team Figure it out together. Managers provide goals, encourage progress, spot issues and remove impediments. Designers, developers and data engineers suggest how to achieve a goal and feed in to continuous improvement. ### 6. Build integrity in Agility needs quality. Automated tests and proven design patterns allow you to focus on smaller parts of the system. A regular flow of insights to act on aids agility. ### 7. Optimise the whole Focus on the entire value stream, not just individual tasks. Align strategy with development. Consider the entire user experience in the design process. ## Three simpler principles If those seem like too many to get started with, I want to introduce three simpler principles that can help you go faster. I came across these in a book about running, which doesn’t seem like the place you’d find inspiration about product management! Think easy, light and smooth. It’s from a man called Micah True who lived in the Mexican desert and went running with the local Native Americans. They called him Caballo Blanco – ‘White Horse’ – because of his speed. > “You start with easy, because if that’s all you get, that’s not so bad. Then work on light. Make it effortless, like you don’t give a shit how high the hill is or how far you’ve got to go. When you’ve practised that so long that you forget you’re practicing, you work on making it smooooooth. You won’t have to worry about the last one – you get those three, and you’ll be fast.” You can do this every cycle. Find one thing to make easier, one thing to make lighter, and one thing to make smoother. Fast will happen naturally.

boringmagi.cc

You don’t have to do fortnightly sprints

In early 2024, we helped GOV.‌UK Design System design and implement a new model for agile delivery. It was a break away from traditional Scrum and two-week sprints towards an emphasis on iteration and reflection. ## Why change things? Traditional two-week sprints and Scrum provide good training wheels for teams who are new to agile, but those don’t work for well established or high performing teams. For research and development work (like discovery and alpha), you need a little bit longer to get your head into a domain and have time to play around making scrappy prototypes. For build work, a two-week sprint isn’t really two weeks. With all the ceremonies required for co-ordination and sharing information – which is a lot more labour-intensive in remote-first settings – you lose a couple of days with two-week sprints. Sprint goals suck too. It’s far too easy to push it along and limp from fortnight to fortnight, never really considering whether you should stop the workstream. It’s better to think about your appetite for doing something, and then to focus on getting valuable iterations out there rather than committing to a whole thing. ## How it works You can see how it works in detail on the GOV.‌UK Design System’s team playbook and in a blog post from the team’s delivery manager, Kelly. There’s also a graphic that brings the four-week cycle to life. There are a few principles that make this method work: * Fixed time, variable scope * Think in iterations: vertical not horizontal slices * Each cycle ends with something shippable or showable * R&D cycles end on decisions around scope * Each cycle starts with a brief, but the team has autonomy over delivery This gives space for ideas and conversations to breathe, for spikes and scrappy prototypes to come together, and for teams to make conscious decisions about scope and delivering value to users. ## How did it work out? In their first cycle, the team delivered three out of five briefs – which was higher than their completion rate at the time. As Kelly reported, ‘most team members enjoyed working in smaller, focused groups and having autonomy over how they deliver their work.’ A few months later, we analysed how often the team was releasing new software: **they were releasing twice as often in half the time.** Between October 2022 and October 2023, there were five releases. Between October 2023 and March 2024, there were 10 releases. One year on and the team has maintained momentum. Iterations have increased, they’ve built a steady rhythm of releasing GOV.‌UK Frontend more frequently, and according to a recent review the team is a lot happier working that way. ## Want to try something new? If you’re looking to increase team happiness and effectiveness, drop us a line and we can chat about transforming your team’s delivery model too.

boringmagi.cc

Our positions on generative AI

Like many trends in technology before it, we’re keeping an eye on artificial intelligence (AI). AI is more of a concept, but generative AI as a general purpose technology has come to the fore due to recent developments in cloud-based computation and machine learning. Plus, technology is more widespread and available to more people, so more people are talking about generative AI – compared to something _even more_ ubiquitous like HTML. Given the hype, it feels worthwhile stating our positions on generative AI – or as we like to call it, ‘applied statistics’. We’re open to working on and with it, but there’s a few ideas we’ll bring to the table. ## The positions 1. Utility trumps hyperbole 2. Augmented not artificial intelligence 3. Local and open first 4. There will be consequences 5. Outcomes over outputs ### Utility trumps hyperbole The fundamental principle to Boring Magic’s work is that people want technologies to work. People prefer things to be functional first; the specific technologies only matter when they reduce or undermine the quality of the utility. There are outsized, unfounded claims being made about the utility of AI. It is not ‘more profound than fire’. The macroeconomic implications of AI are often overstated, but it’ll still likely have an impact on productivity. We think it’s sensible to look at how generative AI can be useful or make things less tedious, so we’re exploring the possibilities: from making analysis more accessible through to automating repeatable tasks. We won’t sell you a bunch of hype, just deliver stuff that works. ### Augmented not artificial intelligence Technologies have an impact on the availability of jobs. The introduction of the digital spreadsheet meant that chartered accountants could easily punch the numbers, leading to accounting clerks becoming surplus to requirements. Jevon’s paradox teaches us that AI will lead to more work, not less. Over time accountants needed fewer clerks, but increases in financial activity have lead to a greater need for auditors. So we will still need people in jobs to do thinking, reasoning, assessing and other things people are good at. Rather than replacing people with machines to reduce costs, technology should be used to empower human workers. We should augment the intelligence of our people, not replace it. That means using things like large language models (LLMs) to reduce the inertia of the blank page problem, helping with brainstorming, rather than asking an LLM to write something for you. Extensive not intensive technology. ### Local and open first Right now, we’re in a hype cycle, with lots of enthusiasm, funding and support for generative AI. The boom of a hype cycle is always followed by a bust, and AI winters have been common for decades. If you add AI to your product or service and rely on a cloud-based supplier for that capability, you could find the supplier goes into administration – or worse, enshittification, when fees go up and the quality of service plunges. And free services are monetised eventually. But there are lots of openly-available generative text and vision models you can run on your own computer – your ‘local machine’ – breaking the reliance on external suppliers. When exploring how to apply generative AI to a client’s problem, we’ll always use an open model and run it locally first. It’s cheaper than using a third party, and it’s more sustainable too. It also mitigates some risks around privacy and security by keeping all data processing local, not running on a machine in a data centre. That means we can get started sooner and do a data protection impact assessment later, when necessary. We can use the big players like OpenAI and Anthropic if we need to, but let’s go local and open first. ### There will be consequences People like to think of technology as a box that does a specific thing, but technology impacts and is impacted by everything around it. Technology exists within an ecology. It’s an inescapable fact, so we should try to sense the likely and unlikely consequences of implementing generative AI – on people, animals, the environment, organisations, policy, society and economies. That sounds like a big project, but there are plenty of tools out there to make it easier. We’ve used tools like consequence scanning, effects mapping, financial forecasting, Four Futures and other extrapolation methods to explore risks and harms in the past. As responsible people, it’s our duty to bring unforeseen consequences more into view, so that we can think about how to mitigate the risks or stop. ### Outcomes over outputs It feels like everyone’s doing something with generative AI at the moment, and, if you’re not, it can lead to feeling left out. But this doesn’t mean you have to do something: FOMO is not a strategy. We’ll take a look at where generative AI might be useful, but we’ll also recommend other technologies if those are cheaper, faster or more sustainable. That might mean implementing search and filtering instead of a chatbot, especially if it’s an interface that more people are used to. It’s more important to get the job done and achieve outcomes, instead of doing the latest thing because it’s cool. ## Let’s be pragmatic Ultimately our approach to generative AI is like any other technology: we’re grounded in practicality, mindful of being responsible and ethical, and will pursue meaningful outcomes. It’s the best way to harness its potential effectively. Beware the AI snake oil.

boringmagi.cc

Metrics, measures and indicators: a few things to bear in mind

Metrics, measures and indicators help you track and evaluate outcomes. They can tell us if we’re moving in the right direction, if things aren’t going well, or if we’ve achieved the outcome we set out to achieve. If you’ve reported on key performance indicators (KPIs), checked progress against objectives and key results (OKRs) or looked at user analytics, you’ll have some experience with metrics, measures and indicators. These words are often used interchangably and, in general, the difference isn’t important. Not for this post anyway. We can talk about the difference between metrics, measures and indicators later. In this post we’ll cover some guiding principles for designing and using metrics, measures and indicators. A few things to bear in mind. ## Guiding principles 1. Value outcomes over outputs 2. Measures, not targets 3. Balance the what (quantitative) and the why (qualitative) 4. Measure the entire product or service 5. Keep it light and actionable 6. Revisit or refine as things change ### Value outcomes over outputs We acknowledge that outputs are on the path to achieving outcomes. You can’t cater for a memorable birthday party without making some sandwiches. But delivering outcomes is the real reason why we’re here. So we don’t measure whether we’ve delivered a product or feature, we measure the impact it’s having. ### Measures, not targets Follow Goodhart’s Law: ‘When a measure becomes a target, it ceases to be a good measure.’ There are numerous factors that contribute to a number or reading going up or down. Metrics, measures and indicators are a starting point for a conversation, so we can ask why and do something about it (or not). The measures are in service of learning: tools, not goals. ### Balance the what (quantitative) and the why (qualitative) Grown-ups love numbers. But it’s very easy to ignore what users think and feel when you only track quantitative measures. Numbers tell us what’s happening, but feedback can tell us why. There’s no point doing something faster if it makes the experience worse for users, for example – we have to balance quantity and quality. ### Measure the entire product or service If we can see where people start, how they move through and where they end, we can identify where to focus our efforts for improvements. The same is true for people who come back too, we want to see whether we’ve made things better than last time they were here. If you’re only measuring one part, you only know how one part is performing. Get holistic readings. ### Keep them light and actionable It’s easy to go overboard and start tracking everything, but too much information can be a bad thing. If we track too many metrics, we run the risk of analysis paralysis. Similarly, one measure is too few: it’s not enough to understand an entire system. Four to eight key metrics or indicators per team is enough and should inspire action. ### Revisit or refine as things change Our priorities will change over time, meaning we will need to change our indicators, measures and metrics too. It’s no use tracking and reporting on datapoints that don’t relate to outcomes. Measure what matters. We should aim not to change them too frequently – that causes whiplash. But it’s all right to change them when you change direction or focus. ## Are we on the way? Or did we get there? Those principles are handy for working out what to measure, but there’s two types of indicator you need to know about: leading and lagging. Leading indicators tell us whether we’re making progress towards an outcome. _Are we on the way?_ For example, if we want to make it easy to find datasets, are people searching for data? Is the number of people searching for data going up? Lagging indicators tell us whether we’ve achieved the outcome. _Did we get there?_ In that same example, making it easy to find datasets, what’s the user satisfaction score? Are they requesting new datasets?

boringmagi.cc

Using quarters as a checkpoint

Breaking your strategy down into smaller, more manageable chunks can help you make more progress sooner. Some things take a long while to achieve, but smaller goals help us celebrate the wins along the way. Many organisations use a quarter – a block of 3 months – to do this. And it can be helpful to look back before you look forward, to celebrate the progress you’ve made and work out what to do next. Every 3 months, we encourage product teams to take the opportunity to step back from the day-to-day and consider the objectives they’re working towards. The quarterly checkpoint is a time to refocus efforts and double-down, change direction or move on to the next objective. There are 2 stages to using the quarterly checkpoint well: 1. Check on your progress 2. Plan how to achieve your new objectives Here are two workshops you can run at each stage, but you can combine them into one workshop if you like. Whatever works. ## Check on your progress First, check on the progress your team has made on your objectives and key results (OKRs). You can do this in a team workshop lasting 30 to 60 minutes. ### 1. List out the OKRs you’ve been working on (10 to 20 mins) Run through the OKRs you’ve been working on. Talk about the progress you made on each key result and celebrate the successes – big or small! ### 2. Think about what’s left to do (20 to 40 mins) For any OKRs you haven’t completed – where progress on key results isn’t 100% – discuss as a team which initiatives you have left to do to fully achieve the objective. For example, you may need to collect some data, run a test, build a thing or achieve an outcome. Consider whether you should change your approach, for example, by doing something smaller or using different methods, based on what you’ve learned over the last quarter. It’s OK to stick to the original plan if it’s still the best approach. Write down what initiatives your team has agreed to do. ## Plan how to achieve your new objectives Next, you’ll need to form a loose plan for how to achieve your new objectives. You can treat unfinished objectives from the previous quarter as a new objective. Run another workshop lasting 30 to 45 minutes for each objective. Everyone on the team will need to input on the plan using the outline below. Write it in a doc, a slide deck or on a whiteboard – whatever works. You will probably want to present these plans to the senior management team or service owner at the start of the new quarter. If it’s easier than starting with a blank page, team leads can fill in the outline and get feedback from the rest of the team. As long as everyone gets a chance to input, it doesn’t matter. It’s OK if you take less than 30 minutes, especially if you already have a plan. ### 1. Write down and describe the objective An objective is a bold and qualitative goal that the organisation wants to achieve. It’s best that they’re ambitious, not super easy to achieve or audacious in nature; they are not sales targets. Write down the problem you’re solving and who it’s a problem for. Discuss how you’ll know when you’re done. What are the success criteria? ### 2. Think about risks and unknowns What might be a challenge? What are the riskiest assumptions or big unknowns to highlight? Do you need to try new techniques? These might form the first initiatives in your plan. You can frame your assumptions using the hypothesis statement: **Because** [of something we think or know] **We believe that** [an initiative or task] **Will achieve** [desired outcome] **Measured by** [metric] Note down dependencies on other teams, for example, where you may need another team to do something for you. ### 3. Detail all the initiatives Write a sentence for all the initiatives – tasks and activities – you’ll need to do to achieve the objective. Consider research and discovery activities, which can help you gather information to turn unknowns into knowns. Consider alphas, things to prototype, spikes, and experiments that can help you de-risk or validate assumptions. Make sure to remember the development and delivery work too – that’s how we release value to users! ### 4. What will you measure? Review your success criteria. Define the metrics that will tell you when you’ve finished or achieved the objective. These should tell you when you’re done and will become your key results. Remember, metrics should be: * tangible and quantitative * specific and measurable * achievable and realistic ### 5. Prioritise radically What would you do differently if you only had half the time? How will you start small and build up? What’s the least amount of work you can do to learn the most? Use these thoughts to consider any changes to your initiatives. Go back and edit the initiatives if you need to. ## Don’t worry about adapting your plans A core tenet of agile is responding to change over following a plan, so don’t be afraid to change your plans based on new information. The quarterly checkpoint isn’t the only time you can look back to look forward – that’s why retrospectives are useful. You can use the activities above at any point. The best product teams build these behaviours into their regular practice. If you’d like help running these workshops or have any questions, get in touch and we’ll set up a chat.

boringmagi.cc

Going faster with Lean principles

Software teams are often asked to go faster. There are many factors that influence the speed at which teams can discover, design and deliver solutions, and those factors aren’t always in a team’s control. But Lean principles offer teams a way to analyse and adapt their operating model – their ways of working. ## What is Lean? Lean is a method of manufacturing that emerged from Toyota’s Production System in the 1950s and 1960s. It’s a system that incorporates methods of production and leadership together. The early Agile community used Lean principles to inspire methods for making digital products and services. These principles have had influence beyond the production environment and have been adapted for business and strategy functions too. ## Books on Lean Four books on Lean principles have influenced the way I work. **1._Lean Software Development: An Agile Toolkit_ by Mary and Tom Poppendieck** The earliest of the four books. It really set the standard. **2._The Lean Startup_ by Eric Ries** This started a big movement for applying Lean principles to your startup, including testing out new business models or growth opportunities. **3._Lean UX_ by Jeff Gothelf and Josh Seiden** One of my favourites. This one really brought strategic goals and user experience closer together. It also shifted teams from writing problem statements to writing hypotheses. **4._The Lean Product Playbook_ by Dan Olsen** This is relatively similar to _The Lean Startup_ but is more of a playbook, showing the practice that goes with the theory. The highlight is its emphasis on MVP tests: experiments you can run to learn something without building anything. ## Lean principles All these books have some principles in their pages, all based on the original Lean principles from Toyota. They’re all pretty similar. Combining their approaches helps us apply Lean principles to business model development, strategy, user-centred design and software delivery. > A note on principles: Principles are not rules. Principles guide your thinking and doing. Rules say what’s right and wrong. ### 1. Eliminate waste Reduce anything which does not help deliver value to the user. So: partially done work; scope creep; re-learning; task-switching; waiting; hand-offs; defects; management activities. Outcomes, not outputs. ### 2. Amplify learning Build, measure, learn. Create feedback loops. Build scrappy prototypes, run spikes. Write tests first. Think in iterations. ### 3. Decide as late as possible Call out the assumptions or uncertainties, try out different options, and make decisions based on facts or evidence. ### 4. Deliver as fast as possible Shorter cycles improve learning and communication, and helps us meet users’ needs as soon as possible. Reduce work in progress, get one thing done, and iterate. ### 5. Empower the team Figure it out together. Managers provide goals, encourage progress, spot issues and remove impediments. Designers, developers and data engineers suggest how to achieve a goal and feed in to continuous improvement. ### 6. Build integrity in Agility needs quality. Automated tests and proven design patterns allow you to focus on smaller parts of the system. A regular flow of insights to act on aids agility. ### 7. Optimise the whole Focus on the entire value stream, not just individual tasks. Align strategy with development. Consider the entire user experience in the design process. ## Three simpler principles If those seem like too many to get started with, I want to introduce three simpler principles that can help you go faster. I came across these in a book about running, which doesn’t seem like the place you’d find inspiration about product management! Think easy, light and smooth. It’s from a man called Micah True who lived in the Mexican desert and went running with the local Native Americans. They called him Caballo Blanco – ‘White Horse’ – because of his speed. > “You start with easy, because if that’s all you get, that’s not so bad. Then work on light. Make it effortless, like you don’t give a shit how high the hill is or how far you’ve got to go. When you’ve practised that so long that you forget you’re practicing, you work on making it smooooooth. You won’t have to worry about the last one – you get those three, and you’ll be fast.” You can do this every cycle. Find one thing to make easier, one thing to make lighter, and one thing to make smoother. Fast will happen naturally.

boringmagi.cc

Our positions on generative AI

Like many trends in technology before it, we’re keeping an eye on artificial intelligence (AI). AI is more of a concept, but generative AI as a general purpose technology has come to the fore due to recent developments in cloud-based computation and machine learning. Plus, technology is more widespread and available to more people, so more people are talking about generative AI – compared to something _even more_ ubiquitous like HTML. Given the hype, it feels worthwhile stating our positions on generative AI – or as we like to call it, ‘applied statistics’. We’re open to working on and with it, but there’s a few ideas we’ll bring to the table. ## The positions 1. Utility trumps hyperbole 2. Augmented not artificial intelligence 3. Local and open first 4. There will be consequences 5. Outcomes over outputs ### Utility trumps hyperbole The fundamental principle to Boring Magic’s work is that people want technologies to work. People prefer things to be functional first; the specific technologies only matter when they reduce or undermine the quality of the utility. There are outsized, unfounded claims being made about the utility of AI. It is not ‘more profound than fire’. The macroeconomic implications of AI are often overstated, but it’ll still likely have an impact on productivity. We think it’s sensible to look at how generative AI can be useful or make things less tedious, so we’re exploring the possibilities: from making analysis more accessible through to automating repeatable tasks. We won’t sell you a bunch of hype, just deliver stuff that works. ### Augmented not artificial intelligence Technologies have an impact on the availability of jobs. The introduction of the digital spreadsheet meant that chartered accountants could easily punch the numbers, leading to accounting clerks becoming surplus to requirements. Jevon’s paradox teaches us that AI will lead to more work, not less. Over time accountants needed fewer clerks, but increases in financial activity have lead to a greater need for auditors. So we will still need people in jobs to do thinking, reasoning, assessing and other things people are good at. Rather than replacing people with machines to reduce costs, technology should be used to empower human workers. We should augment the intelligence of our people, not replace it. That means using things like large language models (LLMs) to reduce the inertia of the blank page problem, helping with brainstorming, rather than asking an LLM to write something for you. Extensive not intensive technology. ### Local and open first Right now, we’re in a hype cycle, with lots of enthusiasm, funding and support for generative AI. The boom of a hype cycle is always followed by a bust, and AI winters have been common for decades. If you add AI to your product or service and rely on a cloud-based supplier for that capability, you could find the supplier goes into administration – or worse, enshittification, when fees go up and the quality of service plunges. And free services are monetised eventually. But there are lots of openly-available generative text and vision models you can run on your own computer – your ‘local machine’ – breaking the reliance on external suppliers. When exploring how to apply generative AI to a client’s problem, we’ll always use an open model and run it locally first. It’s cheaper than using a third party, and it’s more sustainable too. It also mitigates some risks around privacy and security by keeping all data processing local, not running on a machine in a data centre. That means we can get started sooner and do a data protection impact assessment later, when necessary. We can use the big players like OpenAI and Anthropic if we need to, but let’s go local and open first. ### There will be consequences People like to think of technology as a box that does a specific thing, but technology impacts and is impacted by everything around it. Technology exists within an ecology. It’s an inescapable fact, so we should try to sense the likely and unlikely consequences of implementing generative AI – on people, animals, the environment, organisations, policy, society and economies. That sounds like a big project, but there are plenty of tools out there to make it easier. We’ve used tools like consequence scanning, effects mapping, financial forecasting, Four Futures and other extrapolation methods to explore risks and harms in the past. As responsible people, it’s our duty to bring unforeseen consequences more into view, so that we can think about how to mitigate the risks or stop. ### Outcomes over outputs It feels like everyone’s doing something with generative AI at the moment, and, if you’re not, it can lead to feeling left out. But this doesn’t mean you have to do something: FOMO is not a strategy. We’ll take a look at where generative AI might be useful, but we’ll also recommend other technologies if those are cheaper, faster or more sustainable. That might mean implementing search and filtering instead of a chatbot, especially if it’s an interface that more people are used to. It’s more important to get the job done and achieve outcomes, instead of doing the latest thing because it’s cool. ## Let’s be pragmatic Ultimately our approach to generative AI is like any other technology: we’re grounded in practicality, mindful of being responsible and ethical, and will pursue meaningful outcomes. It’s the best way to harness its potential effectively. Beware the AI snake oil.

boringmagi.cc

Tips on doing show & tell well

## What is a show & tell? A show & tell is a regular get-together where people working on a product or service celebrate their work, talk about what they’ve learned, and get feedback from their peers. It’s also a chance to * bring together team members, management and leadership to bond, share success, and collaborate * let colleagues know what you’re working on, keep aligned, and create opportunities to connect and work together * tell stakeholders (including users, partner organisations and leadership) what you’ve been doing and take their questions as feedback (a form of governance). A show & tell may be internal, limited to other people in the same team or organisation, or open to anyone to join. Most teams start with an internal show & tell and make these open later. A show & tell might also be called a team review. ## How to run a great show & tell 1. **Don’t make it up on the spot** Spend time as a team working out what you want to say and who is going to share stories with the audience (1 or 2 people works best). 30 to 60 minutes of prep will pay off. 2. **Set the scene** Always introduce your project or epic. Who’s on the team? What are you working on? What problem are you solving? Who are your users? Why are you doing it? You don’t need to tell the full history, a 30-second overview is enough. 3. **Show the thing!** Scrappy diagrams, Mural boards, Post-it notes, screenshots, scribbles, photos, and clicking through prototypes bring things to life. Text and code is OK, but always aim to demonstrate something working – don’t just talk through a doc or some function. 4. **Talk about what you’ve learned** Share which assumptions turned out to be incorrect, or what facts surprised you. Show clips from user research and usability testing. Highlight important analytics data or performance measures. Share both findings and insights. Be clear on the methodology and any confidence intervals, levels of confidence, risky assumptions, etc. 5. **Be clear** Don’t hide behind jargon. Make bold statements. Say what you actually think! This helps everyone concentrate on the main point, and it generates discussion. 1. **Always share unfinished thinking** Forget about the polish and perfection. A show & tell is the perfect place to collect feedback, ideas and thoughts. It’s a complicated space. We’re all trying to figure it out! 2. **Rehearse** Take 10–15 minutes to rehearse your section with your team to work out whether you need to cut anything. If you’re struggling to edit, use a format like What? So what? Now what? to keep things concise. If you take up more time than you’ve been given, it’ll eat into other people’s section meaning they have to rush (or not share at all) which isn’t fair. 3. **Leave time for questions** The best show & tells have audience participation. Wherever possible, leave time for questions – either after each team or at the end. Encourage people to ask questions in the chat, on Slack, in docs, etc. If you do nothing else, follow tip number 3. You can read more tips on good show & tells from Mark Dalgarno, Emily Webber and Alan Wright. ## How to be a great show & tell audience member 1. **Be present and listen** There’s nothing worse than preparing for a show & tell only to realise that no one’s paying attention. Close Slack, close Teams, stop looking at email, and give your full attention to your team-mates. 2. **Smile, use emojis, and celebrate!** Bring the good vibes and lift each other up whenever there’s something worth celebrating. ## It’s ok to be halfway done The main thing to remember is that show & tell is not just about sharing progress and successes. It’s a time to talk about what’s hard and what didn’t work too. It’s ok to be halfway done. It’s ok to go back to the drawing board. Each sprint, try to answer these questions in your show & tell: * What did we learn or what changed our mind? * What can we show? How can we help people see behind the scenes? * What haven’t we figured out? What do we want feedback on?

boringmagi.cc

You don’t have to do fortnightly sprints

In early 2024, we helped GOV.‌UK Design System design and implement a new model for agile delivery. It was a break away from traditional Scrum and two-week sprints towards an emphasis on iteration and reflection. ## Why change things? Traditional two-week sprints and Scrum provide good training wheels for teams who are new to agile, but those don’t work for well established or high performing teams. For research and development work (like discovery and alpha), you need a little bit longer to get your head into a domain and have time to play around making scrappy prototypes. For build work, a two-week sprint isn’t really two weeks. With all the ceremonies required for co-ordination and sharing information – which is a lot more labour-intensive in remote-first settings – you lose a couple of days with two-week sprints. Sprint goals suck too. It’s far too easy to push it along and limp from fortnight to fortnight, never really considering whether you should stop the workstream. It’s better to think about your appetite for doing something, and then to focus on getting valuable iterations out there rather than committing to a whole thing. ## How it works You can see how it works in detail on the GOV.‌UK Design System’s team playbook and in a blog post from the team’s delivery manager, Kelly. There’s also a graphic that brings the four-week cycle to life. There are a few principles that make this method work: * Fixed time, variable scope * Think in iterations: vertical not horizontal slices * Each cycle ends with something shippable or showable * R&D cycles end on decisions around scope * Each cycle starts with a brief, but the team has autonomy over delivery This gives space for ideas and conversations to breathe, for spikes and scrappy prototypes to come together, and for teams to make conscious decisions about scope and delivering value to users. ## How did it work out? In their first cycle, the team delivered three out of five briefs – which was higher than their completion rate at the time. As Kelly reported, ‘most team members enjoyed working in smaller, focused groups and having autonomy over how they deliver their work.’ A few months later, we analysed how often the team was releasing new software: **they were releasing twice as often in half the time.** Between October 2022 and October 2023, there were five releases. Between October 2023 and March 2024, there were 10 releases. One year on and the team has maintained momentum. Iterations have increased, they’ve built a steady rhythm of releasing GOV.‌UK Frontend more frequently, and according to a recent review the team is a lot happier working that way. ## Want to try something new? If you’re looking to increase team happiness and effectiveness, drop us a line and we can chat about transforming your team’s delivery model too.

boringmagi.cc

Metrics, measures and indicators: a few things to bear in mind

Metrics, measures and indicators help you track and evaluate outcomes. They can tell us if we’re moving in the right direction, if things aren’t going well, or if we’ve achieved the outcome we set out to achieve. If you’ve reported on key performance indicators (KPIs), checked progress against objectives and key results (OKRs) or looked at user analytics, you’ll have some experience with metrics, measures and indicators. These words are often used interchangably and, in general, the difference isn’t important. Not for this post anyway. We can talk about the difference between metrics, measures and indicators later. In this post we’ll cover some guiding principles for designing and using metrics, measures and indicators. A few things to bear in mind. ## Guiding principles 1. Value outcomes over outputs 2. Measures, not targets 3. Balance the what (quantitative) and the why (qualitative) 4. Measure the entire product or service 5. Keep it light and actionable 6. Revisit or refine as things change ### Value outcomes over outputs We acknowledge that outputs are on the path to achieving outcomes. You can’t cater for a memorable birthday party without making some sandwiches. But delivering outcomes is the real reason why we’re here. So we don’t measure whether we’ve delivered a product or feature, we measure the impact it’s having. ### Measures, not targets Follow Goodhart’s Law: ‘When a measure becomes a target, it ceases to be a good measure.’ There are numerous factors that contribute to a number or reading going up or down. Metrics, measures and indicators are a starting point for a conversation, so we can ask why and do something about it (or not). The measures are in service of learning: tools, not goals. ### Balance the what (quantitative) and the why (qualitative) Grown-ups love numbers. But it’s very easy to ignore what users think and feel when you only track quantitative measures. Numbers tell us what’s happening, but feedback can tell us why. There’s no point doing something faster if it makes the experience worse for users, for example – we have to balance quantity and quality. ### Measure the entire product or service If we can see where people start, how they move through and where they end, we can identify where to focus our efforts for improvements. The same is true for people who come back too, we want to see whether we’ve made things better than last time they were here. If you’re only measuring one part, you only know how one part is performing. Get holistic readings. ### Keep them light and actionable It’s easy to go overboard and start tracking everything, but too much information can be a bad thing. If we track too many metrics, we run the risk of analysis paralysis. Similarly, one measure is too few: it’s not enough to understand an entire system. Four to eight key metrics or indicators per team is enough and should inspire action. ### Revisit or refine as things change Our priorities will change over time, meaning we will need to change our indicators, measures and metrics too. It’s no use tracking and reporting on datapoints that don’t relate to outcomes. Measure what matters. We should aim not to change them too frequently – that causes whiplash. But it’s all right to change them when you change direction or focus. ## Are we on the way? Or did we get there? Those principles are handy for working out what to measure, but there’s two types of indicator you need to know about: leading and lagging. Leading indicators tell us whether we’re making progress towards an outcome. _Are we on the way?_ For example, if we want to make it easy to find datasets, are people searching for data? Is the number of people searching for data going up? Lagging indicators tell us whether we’ve achieved the outcome. _Did we get there?_ In that same example, making it easy to find datasets, what’s the user satisfaction score? Are they requesting new datasets?

boringmagi.cc