Data-informed design is the practice of systematically collecting, interpreting, and acting on evidence about learner behavior and outcomes throughout the design, development, and deployment of learning solutions. It is not simply the use of data to evaluate a finished product; rather, data shapes design decisions at every stage of the iterative cycle. This approach distinguishes learning engineering from conventional instructional design by making the empirical test of design hypotheses — rather than expert judgment alone — the primary driver of revision. The Learning Engineering Toolkit frames this as one of the field's defining commitments: the learning engineer is expected to identify measurable outcomes, instrument the learning environment to observe them, and revise designs in response to what the data shows. [LE-LS-GL-007]
From Educational Data Mining to Design Practice
The scientific infrastructure that makes data-informed design feasible was established through the development of educational data mining (EDM) as a formal discipline. Ryan Baker and Kalina Yacef's foundational review codified EDM's methods — prediction, clustering, knowledge discovery, relationship mining, and distillation for human judgment — and positioned them as applicable to the analysis of learning system telemetry at scale. [LE-LS-AP-005] These methods allow learning engineers to move from raw interaction logs to actionable design findings: which sequences produce learning gain, where learners disengage, which assessment items are misaligned with instruction, and how different learner populations respond differently to the same design.
The Learning Engineering for Online Education volume documented how digital learning platforms make continuous telemetry collection practical and economical in ways that were not possible in face-to-face settings. [LE-LS-GL-008] Kenneth Koedinger's founding of DataShop — the world's largest open repository of educational log data — created a shared infrastructure that enabled researchers and practitioners to develop and validate data analysis methods on large, authentic datasets rather than small laboratory samples. [LE-LS-PP-005] The Penn Center for Learning Analytics, led by Ryan Baker, extended this infrastructure into policy-relevant research, including the landmark "High-Leverage Opportunities for Learning Engineering" report that identified shared data infrastructure as one of the field's most critical needs. [LE-LS-CO-004]
Data in the Design Cycle
Data-informed design operates through repeated cycles that connect measurement to revision. At the outset of a design cycle, the team specifies what learning outcomes are targeted and how they will be measured — establishing the empirical criteria against which design success will be judged. During deployment, interaction data (time on task, error patterns, hint requests, skip rates, assessment performance) is collected and analyzed to characterize learner behavior. These findings then drive specific, documented revisions to content, sequence, assessment, or interface — with the rationale for each change recorded so that future iterations can evaluate whether the change produced the expected effect. [LE-LS-GL-007]
The practical power of this cycle is illustrated by large-scale evidence on the doer effect: analysis of platform telemetry across seven online courses demonstrated that active practice (doing) causally improves learning outcomes relative to passive reading, and that the magnitude of this effect is consistent across diverse course contexts. [LE-LS-AP-012] This finding, made possible by data-rich platform instrumentation, gives learning engineers an evidence-based design principle — prioritize active practice over passive consumption — that is directly actionable within iterative design cycles. Bayesian Knowledge Tracing, the probabilistic student model developed by Corbett and Anderson that estimates per-student per-skill mastery from response data, is another foundational technology that operationalizes data-informed design in real-time adaptive systems. [LE-LS-AP-011]
¶ Subsections:
From the Learning Engineering Toolkit
The Learning Engineering Toolkit holds that making decisions from data is indispensable to the field, putting it bluntly: leave data out and what remains may be learning design, but it is no longer learning engineering. [LET-05] Data-informed decision-making breaks into two components, instrumentation and analytics. [LET-05] Instrumentation covers designing, building, and deploying the data collection embedded in a learning solution so that it can guide iterative refinement; analytics covers making sense of those data and putting them to use toward the same goal. [LET-05] Instrumentation provides the sensors and data pipelines that record what learners do, while analytics — the disciplined computational study of that record — converts it into findings that steer design. [LET-05], [LET-06]
A defining trait of the learning engineering process is that it draws on data throughout, weighing evidence at every stage rather than saving it for the end. [LET-01] Deciding what data to instrument and defining it belong to understanding the challenge; the instrumentation itself is designed and constructed during creation, exercised during implementation, and examined during investigation to surface ways to improve. [LET-05] This sets learning engineering apart from building a course first and only afterward inviting researchers to gather data for an efficacy study — from the start, the process expects data to shape at least one design iteration. [LET-05]
Working from data matters because people's hunches about what aids learning frequently miss the mark, and examining results from lightweight trials keeps effort trained on the design features that matter most and steers teams away from expensive dead ends. [LET-06] In one case a Kaplan team weighing ways to prepare students for LSAT logical reasoning formed a hypothesis, built an alternative, tried it out in a real setting, and let the evidence settle which option was best — discovering that a short pass through worked examples beat a ninety-minute video. [LET-06] Comparisons like this can be scaled up: the Toolkit observes that teams can run A/B tests of competing design features and iteratively probe which conditions serve particular learners better. [LET-06]
Because engineering is a matter of trade-offs, amassing more data is not automatically an improvement, and teams have to judge which portion of the data bears on the challenge at hand. [LET-05] Research of this sort yields findings tightly bound to specific products and learner groups rather than universal laws, which is why the data cycle has to be run afresh as designs and settings shift. [LET-06]
Sources from the Learning Engineering Toolkit
- [LET-01]Aaron Kessler, Scotty D. Craig, Jim Goodell, Dina Kurzweil & Scott W. Greenwald (2023). Chapter 1: Learning Engineering is a Process. In Jim Goodell & Janet Kolodner, Learning Engineering Toolkit (pp. 29–46). Routledge / Taylor & Francis. Open Access
- [LET-05]Erin Czerwinski, Jim Goodell, Steve Ritter, Robert Sottilare, Khanh-Phuong Thai & Daniel Jacobs (2023). Chapter 5: Learning Engineering Uses Data (Part 1): Instrumentation. In Jim Goodell & Janet Kolodner, Learning Engineering Toolkit (pp. 153–173). Routledge / Taylor & Francis. doi:10.4324/9781003276579
- [LET-06]Michelle Barrett, Erin Czerwinski, Jim Goodell, Daniel Jacobs, Steve Ritter, Robert Sottilare & Khanh-Phuong Thai (2023). Chapter 6: Learning Engineering Uses Data (Part 2): Analytics. In Jim Goodell & Janet Kolodner, Learning Engineering Toolkit (pp. 175–199). Routledge / Taylor & Francis. doi:10.4324/9781003276579