Confessions of a Former District Data Scientist: Why Most K–12 Programs Go Unevaluated

During the first day of the 2026 AI4MN Thought Leaders Summit, I heard something that seemed surprising to some but is common knowledge to me and my colleagues. Chris Agnew, Director of Stanford’s AI for Education Hub, shared two uncomfortable truths about the current state of K–12 education.

First, education has very little high-quality causal research showing the impact of student-facing artificial intelligence. Second, education leaders cannot afford to wait for perfect evidence before making decisions.

The evidence gap may be especially visible with AI, but it did not begin with AI.

For years, school district leaders have made consequential decisions about educational products, services, interventions, curricula, and instructional initiatives with limited local evidence. They still have students to teach, budgets to manage, contracts to renew, and programs to improve. They cannot simply stop making decisions until perfect research arrives.

The K–12 Evidence Problem Is Bigger Than AI

Large school districts may implement dozens or even hundreds of initiatives across their schools. These can include tutoring programs, behavioral interventions, professional development, new curricula, educational technology, after-school services, and nonprofit partnerships.

Only a small portion of them receive a rigorous education program evaluation.

Some initiatives may have research supporting them, but that research usually comes from other districts serving different students under different conditions. Student demographics, staffing, baseline achievement, instructional priorities, access, implementation, and local policies can all influence results.

National research can answer an important question: Has this program worked somewhere? Local evaluation answers a different one: Is this program improving outcomes for our students? Both questions matter, but they are not interchangeable.

I Saw the Problem From Inside a Large School District

I saw this firsthand while working as a data scientist in the Research and Evaluation department at Minneapolis Public Schools.

I was fortunate to work in a department of 13 people. Nearly everyone held a master’s degree or a Ph.D., and the team had the expertise to conduct rigorous research and analyze student data for schools.

Despite that expertise, most district initiatives still went unevaluated. It was not because the team was indifferent. It was not because the district lacked talented people. It was because the department was already responsible for an enormous amount of essential work.

There were compliance reports, assessment analyses, data-cleaning projects, surveys, grant reports, school-board questions, recurring reports, external data requests, and urgent requests from district leadership. New priorities could emerge at any time, often displacing longer-term evaluation projects.

Most initiatives did not go unevaluated because no one cared. They went unevaluated because the people capable of answering the questions were overwhelmed by other necessary work.

This pattern plays out in districts across the country. Internal data teams are invaluable, but they cannot evaluate everything while also meeting every operational, reporting, and strategic demand placed on them.

Why Rigorous Evaluation Falls to the Bottom of the List

Traditional program evaluation for school districts can require significant coordination.

A conventional study may involve designing a research plan, recruiting participants, establishing agreements, collecting new data, coordinating with schools, monitoring implementation, waiting for outcomes, conducting the analysis, and producing a lengthy technical report.

That process may be appropriate for some initiatives. It is not practical for every purchasing, renewal, or expansion decision a district faces.

Timing creates another problem. District leaders may need to decide whether to renew a software contract, continue a nonprofit partnership, expand a pilot, or redirect funds within the current budget cycle. An evaluation that produces results after the decision has already been made offers limited practical value.

Districts need credible evidence, but the traditional process for producing it often conflicts with the timelines and staffing realities of district decision-making.

Consulting From the Outside Did Not Immediately Solve It

After leaving Minneapolis Public Schools, I began providing educational program evaluation consulting to EdTech companies and educational nonprofits.

These organizations hired me to partner with districts, gather relevant data, and estimate the impact of their products or services on student outcomes. The work was valuable, but it was also lengthy, expensive, and highly customized.

Sometimes, after investing considerable time and money, an organization learned that students who participated in its program performed no better than similar students who did not.

That does not mean the evaluation failed. A null or negative finding can prevent an organization from scaling an ineffective approach. It can help a product team identify areas for improvement. Most importantly, it can protect students from spending additional time on something that is not helping them.

The greater failure is allowing a program to continue for years without asking what difference it is making. When evaluation is too difficult to conduct, ineffective programs can survive, effective programs can remain underfunded, and students bear the consequences of both errors.

A More Practical Model for Impact Evaluation

I began asking a different question. Instead of trying to make districts fit the traditional evaluation process, what would it look like to build an evaluation process around the realities districts already face?

The solution needed to meet three requirements.

First, it had to be easy for district staff. That meant using existing district data whenever possible, limiting meetings, and avoiding unnecessary surveys or new collection requirements.

Second, it had to produce results while decisions were still being made. Evidence should arrive in time to inform budget planning, contract renewals, program improvements, and expansion decisions.

Third, it still had to be rigorous. A useful rapid-cycle evaluation must examine real student outcome metrics, compare participants with similar nonparticipants, account for important preexisting differences, and report the results honestly.

Rapid does not mean careless. Speed should come from eliminating unnecessary burden and making focused use of existing data, not from lowering the standard of evidence.

Building MomentMN Snapshot Reports

I began developing this approach in 2019 in partnership with Hopkins Public Schools, my alma mater. After several years of refinement, Parsimony officially launched MomentMN Snapshot Reports in late 2025.

Instead of requiring districts to launch another major research project, MomentMN was designed to work with data districts already collect every day. It uses existing district data and a retrospective, quasi-experimental design. MomentMN uses pre- and post-program outcomes together with multivariate matching to compare students who received a product or service with a carefully constructed group of students who were similar before the initiative began.

The findings are then translated into a clear report and personalized interpretation designed to inform real-world decisions rather than sit unread on a shelf. This approach offers independent program evaluation for K–12 organizations without placing another major project on district staff.

The Question Districts Can No Longer Leave Unanswered

Over the years, I have attended educational advisory meetings as an evaluator, community member, and parent. The people around the table may include district leaders, educators, families, community representatives, and public officials.

Eventually, someone asks the same question: How do we know this is working?

Too often, the response includes participation totals, implementation updates, satisfaction surveys, testimonials, or usage metrics. Those measures may be informative, but they do not show whether student outcomes improved.

That question does not have to remain unanswered simply because a district lacks the time or staff for a multiyear study.

Districts do not need to evaluate everything at once. They need a repeatable, low-burden way to evaluate the initiatives connected to their most consequential decisions.

See What a MomentMN Snapshot Report Delivers

To experience what it is like to receive clear, independent evidence about the impact of you initiatives on the students in your district, email us at [email protected].

We will send you a sample MomentMN Snapshot Report so you can see how existing student data can be transformed into timely, decision-ready findings.

Share

Continue Reading

Experience an Easier Way To Get Rigorous Evidence of Your Impact

Have questions? Want a demo?

Book a call with Dr. Amanuel Medhanie, who will answer your questions and show you
around the Snapshot Report service.