Thought LeadershipBlog

From Policy to Practice: Using Applied Evidence in Education Sandboxes

read

|

30 Mar 2026

|

This is part of a series following our sandbox journey across Southeast Asia — read our first blog here.

As Mike Tyson famously said, “Everyone has a plan until they get punched in the face.” 

In education policy and programmes, this rings uncomfortably true. Well-designed strategies that look strong on paper can meet very different realities once they interact with real classrooms, communities, and systems.

This isn’t a reason to give up on planning; it’s a reason to build in structured ways to test, learn, and adapt. That’s the core premise of the Sandbox Method we’ve been applying across Southeast Asia. We’ve worked with Ministries of Education in Indonesia and the Philippines on policy development, and with SEAMEO VOCTECH on technical and vocational education across the region. Sandboxes create spaces to run real-world experiments, gather feedback, and refine ideas before wider rollout.

But testing ideas requires evidence, and this is where things often get complicated.

When people hear ‘evidence’, they often think of large-scale randomised controlled trials (RCTs), systematic reviews, or multi-year evaluations. These approaches have their place. But they are rarely practical — or even appropriate — in early, iterative phases of policy or programme design, where uncertainty is high, and there isn’t always one ‘right’ answer. 

What we need instead is a different kind of rigour: applied evidence that is fast enough to be useful, targeted enough to answer the right question at the right time, and good enough to move us forward.

Our blog is about what that looks like in practice.

Start With What You Believe Needs to Be True for Your Idea to Succeed and Test That

Every intervention or policy rests on a set of beliefs: assumptions about how change will happen. In sandboxes, we make these assumptions explicit and treat them as hypotheses to be tested rather than facts to rely on.

A useful starting point is asking: what needs to be true for this idea to work? These are your critical beliefs: they are the things most vital to your success, but you have the least information about.

Some will be about value (is there demand for this?), others about impact (does it work?), and others about growth (will it scale and be sustainable?). 

Getting clear on which assumptions carry the most risk — the ones that, if wrong, would derail everything — tells you where to focus your evidence-gathering first.

Chang, A. M. (2018). Lean impact: How to innovate for radically greater social good.

This is the Lean Impact principle that Sandboxes are built on: not just as a one-off check, but as a live lens for deciding what kinds of evidence you need and when.

Key questions to ask at this stage:

  • What are the core assumptions underpinning this intervention or policy?
  • Which of these, if wrong, would most threaten the success of the programme?
  • Is this a question of value (is there demand for this?), impact (does it work?), or growth (will it scale and be sustainable?). 
  • What is the smallest, fastest piece of evidence that would allow us to validate or challenge this assumption?

Example from SEAMEO VOCTECH: When we began our sandbox collaboration with SEAMEO VOCTECH on their SEA-VET Learning platform, one of the first things we looked at was platform analytics. Low engagement metrics — like minutes spent on a course — seemed to confirm an assumption about low user interest. But through early conversations with the SEAMEO VOCTECH team, we quickly learned that learners often download content and engage offline, meaning web analytics alone were a poor proxy for real engagement. Testing this assumption early, using a simple metric review and team discussion rather than a full study, prevented us from designing the wrong intervention entirely.

Collect Just Enough Evidence That You Can Act On — No More, No Less

One of the most important mindset shifts in sandbox thinking is around what counts as enough evidence. The Sandbox Method is deliberately different from academic research or large-scale programme evaluations. We’re not trying to produce findings that are generalisable across contexts. We’re trying to help a specific team make a specific decision about their specific programme and move one step forward.

This is what we mean by right-sizing evidence. The goal isn’t to gather the most data, but to gather the right data, at the right time, in the right amount.

So, what determines ‘just enough’?

A useful frame here is to ask: Does this evidence give us enough confidence to make the next decision? Not to conclude, but to proceed. If you’re testing whether a tool is comprehensible to users, you don’t need a survey of 500 schools. You might need structured observations with 10 teachers and a short debrief. If you’re checking whether a policy workshop format creates genuine engagement, a few well-facilitated sessions with reflective documentation might tell you what you need to know.

The Getting Rigour Right framework and Nesta’s Standards of Evidence both offer useful ways of thinking about what level of evidence is appropriate at different stages. The key insight from these frameworks is that rigour is contextual. What’s rigorous in a Do-Measure-Learn cycle, a fixed, time-boxed period—typically 1 to 4 weeks—during which a team completes a set amount of work, is different from what’s rigorous in a national evaluation. Neither is wrong; they serve different purposes.

Synowiec, C., Fletcher, E., Heinkel, L., & Salisbury, T. (2023). Getting rigor right: A framework for methodological choice in adaptive monitoring and evaluation. Global Health: Science and Practice, 11(Supplement 2).

Example from MoPSE (Indonesia): Our initial assumption in the Ministry of Primary and Secondary Education (MoPSE) sandbox was that policy teams lacked standardised processes and needed toolkits to better integrate evidence into policy. But early interviews told a different story: processes did exist, and the real barrier was documentation, leadership buy-in and culture. The evidence that shifted our thinking wasn’t a large-scale survey; it was a handful of targeted interviews with technical staff, supplemented by process documentation. That combination gave us enough to change direction. The lesson: depth over scale and combining sources for a fuller picture.

Match the Method to the Question, Not the Other Way Round

One of the most common mistakes in implementation research is choosing a method to test and gather evidence and then fitting the question to it, rather than starting with the question and choosing the method that best answers it. Sandboxes require a broader, more pragmatic toolkit.

Different types of questions call for different types of experimentation and evidence generation:

  • Value questions (do people want this? Does it resonate?) → user interviews, observation, feedback sessions, ‘fake door’ tests
  • Impact questions (can this be delivered? Does it work?) → process documentation, workflow walkthroughs, piloting with a small group
  • Growth questions (is this sustainable? does it fit the system?) → stakeholder interviews, co-design workshops, review of institutional context

The ‘fake door’ test, for instance, is a lightweight way to test desirability before building anything: you present a prototype or concept to potential users and observe whether they engage with it as intended. We used this approach with SEAMEO VOCTECH to test assumptions about course format preferences before investing in redesign — gathering real user signals without a full build.

Meanwhile, workshops and facilitated discussions — such as those we ran with Ministries of Education in the Philippines and Indonesia, and with teachers in schools — can quickly surface rich, contextual evidence. They’re not anecdotal: if designed well, with clear prompts, documentation, and deliberate reflection, they produce systematic insight. The key is being intentional about what you’re trying to learn and how you’ll capture it.

The question to ask yourself

What method gives me the clearest signal on this specific assumption with the resources I actually have?

The People Closest to the Problem Have the Best Insights

Applied evidence isn’t just about methods — it’s about who generates it. In our sandboxes, some of the most important insights came not from external researchers or senior leaders, but from the people doing the work every day.

In the DepEd (Philippines) sandbox, focus group discussions with school heads and teachers revealed how a policy tool to assess digital maturity in schools was actually perceived and used at the school level. Their feedback didn’t just flag surface issues; it led us to revisit the policy development process itself. Self-reported data, it turned out, was being shaped by school-level context: infrastructure gaps, resource constraints, and variation in digital competency meant that lower scores reflected structural realities rather than lack of effort.

This kind of insight can’t be generated by analytics alone. It requires bringing the right voices into the room, and often, going to where those voices are.

User-centred approaches to evidence gathering — drawing on participatory research methods and co-design — recognise that the people experiencing a policy or programme are also the best sources of evidence about whether it’s working. This isn’t soft evidence. Structured well, it’s some of the most actionable evidence you can get.

Launch Is the Starting Point, Not the Finish Line

Evidence collection shouldn’t stop when a policy launches or a programme rolls out. One of the biggest gaps in implementation research is what happens after go-live — when assumptions meet reality at scale.

In our policy sandbox work with the Ministry of Education in Indonesia, we built a toolkit that advocates for flex loops and evidence windows: structured moments to pause, review what’s happening, and decide whether to adjust course. 

These aren’t bureaucratic check-ins. They’re designed with specific questions in mind: are the assumptions we launched with still holding? Are we seeing what we expected? What do we need to investigate further?

This requires, as one implementation research review puts it, a shift from verifying what works to continuously adapting it: from avoiding policy failure to assuming it will happen and setting up to learn from it.

The SEAMEO VOCTECH sandbox is a good example of this in action. Rather than treating each design as a final product, the team built in review cycles to look at what was working, what wasn’t, and what could be improved. The sandbox became a mechanism not just for testing ideas but for continuous improvement, with evidence driving the iteration, not instinct alone.

Putting it all together: a quick guide to applied evidence in sandboxes

Ask these questions at the start of each Sandbox sprint:

  • What are we assuming is true, and which assumptions are riskiest?
  • Is this a value, impact, or growth question?
  • What’s the smallest, fastest evidence we need to make the next decision?
  • Who is closest to this problem, and how do we bring them into the evidence process?
  • Have we built in a review window to check assumptions after launch?

Evidence Is Not a Destination, It’s a Practice

What runs through all of this is a shift in how we think about evidence in policy and programme work. It’s not something you gather once, at the end, to prove that something worked. It’s a continuous practice of testing beliefs, surfacing surprises, and adjusting course.

The sandboxes we’re running with Ministries of Education in Indonesia and the Philippines, and SEAMEO VOCTECH, are not just about finding the right answer to a policy challenge. They’re about building the habit and capacity to keep asking the right questions and collecting the right evidence to answer them.

That, more than any single finding at any single moment, is what helps education systems get better over time.

This work is part of the portfolio of projects delivered by the ASEAN-UK SAGE programme. The ASEAN-UK SAGE programme is delivered by the British Council and SEAMEO Secretariat, in partnership with EdTech Hub and Australian Council for Educational Research (ACER), and is an ASEAN cooperation programme funded by the UK.

Share: