
EdTech Software Development Cost in 2026: How to Budget for Your BuildRead More

An adaptive learning platform makes two kinds of decisions: what a learner should do next, and what should happen inside the task they are working on right now. Every layer in the architecture exists to serve one of those two decisions. The order those layers get built in decides whether completion rates move. This guide covers the six layers, what each one needs from your data and content, and where a build should start.
Kurt VanLehn separated tutoring systems into an outer loop and an inner loop. The outer loop runs once per task and picks the next task. The inner loop runs once per step the learner takes inside a task, gives feedback and hints on that step, and updates the learner model that the outer loop then reads. (Source: The Behavior of Tutoring Systems - ACM Digital Library)
The two loops cost different things. Outer-loop decisions happen at module boundaries, a few times an hour, and they can afford a few hundred milliseconds of thinking. Inner-loop decisions happen on every answer, every hint request and every long pause, inside a screen the learner is staring at. An architecture that treats both as one recommendation service ends up too slow for the inner loop or too crude for the outer one.
Adaptive personalization works often enough to justify building and unreliably enough to justify measuring. A 2024 scoping review in Heliyon examined 69 studies of personalized adaptive learning in higher education: 41 of them (59%) reported increased academic performance, and 28 (41%) reported no significant impact. Twenty-five studies (36%) reported increased engagement, and 64% did not report engagement as an outcome.
The tutoring literature sets a ceiling worth knowing before anyone promises a board a step change. VanLehn's 2011 meta-review of 28 evaluation studies found human tutoring raised test scores by an effect size of 0.79 against no tutoring, and step-based computer tutoring reached 0.76. Both fall well short of the two standard deviations often attributed to Benjamin Bloom's 1984 paper, a figure later reviews have not reproduced. A 2024 meta-analysis of 27 randomized experiments on personalized and adaptive learning technologies for K-12 reading reported a pooled effect of g = 0.29.
Careful designs do better. A randomized controlled trial published in Scientific Reports in June 2025 ran a GPT-4 tutor against an active-learning physics class at Harvard with 194 students, and measured an effect size of 0.63 by linear regression, with a median 49 minutes of time on task. The authors are specific about the conditions that produced it: expert-authored prompts, complete step-by-step solutions supplied to the model to limit hallucination, scaffolding built into the platform instead of the prompt, and lessons aimed at the understand, apply and analyze levels of Bloom's taxonomy. They do not claim the result holds for higher-order synthesis.
The Heliyon reviewers attribute the spread to context, writing that effectiveness "may depend on various factors, including the specific implementation strategy, subject matter, and context." That is the argument for everything below. The outcome lives in the implementation.
An adaptive learning platform has six layers, and each one exists to move a different number.
| Layer | The decision it makes | What it moves |
|---|---|---|
| Content and competency model | Which items exist, what each one teaches, what it depends on | Whether personalization is possible at all |
| Instrumentation and event pipeline | What the system knows about a learner, and how quickly it knows it | Every downstream metric, because it caps all of them |
| Learner model | How much this learner knows right now, per competency | Routing accuracy, so time to competency |
| Decision layer | What to serve next, and what to do inside the current task | Completion and drop-off |
| Delivery | How a decision reaches the learner, and what happens when it is slow | Session length and engagement |
| Measurement and experimentation | Whether any of the above worked | Retention, and your ability to defend the roadmap |
Read that as a dependency chain. A learner model trained on events you never captured is not a modelling problem, and a decision layer routing through content that was never broken into routable pieces has nothing to route.
A platform can adapt only at the granularity its content is broken into. A 40-minute video and a 12-question end-of-module quiz give the decision layer two available moves: play it, or skip it.
Three pieces of work turn a course library into something a system can route through. Atomization breaks courses into items small enough to serve individually and specific enough to assess: a worked example, a single concept explanation, one question. The competency graph names what each item teaches and which competencies depend on which, giving the decision layer edges to traverse. Item metadata records difficulty, discrimination, expected time and media type, so a routing decision can account for more than the topic.
Assessment items are what most working systems actually run on. In the Heliyon review, pre-knowledge quizzes were the most common trigger for adaptive content delivery at 58%, ahead of learning analytics and activity logs at 22%. That makes the item bank the gate on everything after it.
Calibration is the part teams underestimate. Difficulty and discrimination get estimated from real learner responses, which means the first cohort through a new item is doing calibration work whether or not anyone planned for it. On how much data that takes, a 2025 tutorial in Advances in Methods and Practices in Psychological Sciencerecommends Monte Carlo simulation to plan sample size in advance, because the answer depends on the model, the test design and which parameters matter to you. Published rules of thumb vary widely enough that running the simulation against your own design beats copying a number.
Give the content workstream its own owner and its own schedule. Authoring and tagging have no shortcut, and every later phase depends on them, which makes this the part of the plan most likely to move other dates.
Capture every learner interaction as a timestamped event scoped to a learner and an item, in a schema that will survive a model change.
xAPI, maintained by the ADL Initiative at the US Department of Defense, is the usual choice. Every xAPI statement carries an actor, a verb and an object, and statements go to a Learning Record Store, which the specification defines as "a server ... responsible for receiving, storing, and providing access to Learning Records." Caliper Analytics from 1EdTech covers similar ground with a different shape, using a Sensor API to transmit event data and metric profiles to give each kind of learning activity a shared vocabulary.
Open edX published its reasoning, which makes it a useful worked example. The project adopted xAPI as the default standard for its analytics system and converts legacy tracking logs into xAPI statements through an event-routing-backends package before they reach an LRS. Caliper was set aside for a plain reason: the community asked for xAPI repeatedly and never requested Caliper. The same record names a requirement that is easy to skip, a shared and versioned transformation library serving both the real-time and the historical pipeline, so that a change to an event's shape cannot make last month's data mean something different from this month's.
One rule from Google's machine learning guidance belongs here rather than in the modelling section. Rule 29 reads: "The best way to make sure that you train like you serve is to save the set of features used at serving time, and then pipe those features to a log to use them at training time." A learner model trained on features reconstructed after the fact behaves differently in production than it did in evaluation, and that gap is painful to diagnose six months later.
Real time in an adaptive learning platform means three separate latency budgets. Jakob Nielsen's response-time limits still set the targets: 0.1 seconds feels instantaneous, 1 second keeps a person's flow of thought uninterrupted, and 10 seconds is the outer limit of held attention.
Decision | Loop | Budget | Where it runs |
|---|---|---|---|
Feedback on an answer | Inner | 0.1 to 1 second | In the request, precomputed or rule-based |
Hint selection inside a task | Inner | Under 1 second | In the request, on cached features |
Next item in the session | Outer | Under 1 second end to end, so inference gets a fraction of it | Online inference against a feature store |
Revision to the learner's plan | Outer | Seconds to minutes | Async, after the interaction |
Item recalibration and model retraining | Neither | Hours | Nightly batch |
Every online inference call needs a timeout and a rule-based answer waiting behind it. A learner who waits four seconds for a personalized next question has had a worse session than one who got a fixed sequence instantly.
Connection quality belongs in the same budget. Arbisoft built Philanthropy University on Open edX with NodeBB for its community side, with web and mobile apps designed to work over low-bandwidth and intermittent connections, for a platform the case study says has over 100,000 registered users. Under those conditions an adaptive path that needs a server round trip per step stops working, and the design has to push a short queue of already-decided items to the client.
Start with rules and a mastery threshold, then earn the right to anything more complex. Google's Rule 1 is "Don't be afraid to launch a product without machine learning" and Rule 4 is "Keep the first model simple and get the infrastructure right." Both land hard here, because the rules baseline doubles as the control condition every later model has to beat.
Rules and thresholds mark a competency as mastered at N correct out of M recent attempts and block downstream content behind unmastered prerequisites. The logic is explainable to a learner and to a procurement committee, and it is cheap enough to sit in the request path without a latency discussion.
Item response theory places item difficulty and learner ability on one scale, which is what assessment-heavy platforms and certification products need for defensibility. It assumes ability holds still during measurement, so it models testing better than it models learning.
Bayesian knowledge tracing runs a hidden Markov model per competency and updates the probability that a learner has mastered it after each response. It handles change over time, which item response theory does not, and it stays interpretable.
Deep knowledge tracing replaces that with a recurrent network. Piech and colleagues' 2015 paper reported an LSTM reaching AUC 0.85 on 1.4 million Khan Academy answers from 47,495 students, against 0.68 for standard Bayesian knowledge tracing, and 0.86 on the Assistments benchmark against a best prior result of 0.69.
Those numbers need a correction that rarely travels with them. The 2022 pyKT benchmark re-ran ten deep knowledge tracing models across seven datasets under one preprocessing and evaluation protocol. Reported AUC "of the same approach on the same dataset vary surprisingly from 0.709 to 0.86" between papers, and "wrong evaluation setting may cause label leakage that generally leads to performance inflation." The authors also found that "the improvement of many DLKT approaches is minimal compared to the very first DLKT model proposed by Piech et al." A deep model is a reasonable thing to reach for once you hold sequence data at that volume and an evaluation protocol you trust. Reaching for it first buys a number nobody can interpret.
Contextual bandits fit the case where several plausible content variants exist and no knowledge model distinguishes them. They also suit cold start, since exploration is the point rather than a side effect.
Large language models are the strongest inner-loop component available and a weak outer loop on their own. The Harvard trial's design notes are more instructive than its effect size: the model was handed expert-authored prompts and complete step-by-step solutions, and the sequencing lived in the platform. A language model has no calibrated estimate of what a learner knows across a competency graph and no auditable record of how a routing decision got made, which is what the outer loop is accountable for.
Cold start has two workable answers, and both depend on the item bank. Run a short pre-knowledge quiz, which is what most studied systems do, or start from population priors for similar learners and let a bandit explore for the first few items.
Pick the unit of randomization and the guardrail metrics before the first model ships, because neither can be chosen honestly afterwards.
Randomize learners, not sessions. Completion is a per-learner outcome, and a learner who experiences both conditions contaminates both arms. Keep a holdout cohort on the non-adaptive path permanently rather than for one experiment. An adaptive system changes what content people see, which changes what "the same course" means over time, and a year in there is no baseline left to compare against.
Guardrails matter as much as the headline number. Track time to first success, the point in each module where learners leave, remediation loops per competency, help-seeking rate, and assessment score on a fixed non-adaptive instrument. The last one catches the most common self-deception, where completion rises because the system routed learners around the hard material.
Watch for training-serving skew, Google's Rule 37, which is the gap between how a model performs on held-out data and how it performs on live traffic. The usual cause is a feature computed one way in the training pipeline and another way in the serving path. Logging features at serving time, per Rule 29, is what makes the gap visible.
Recall that 64% of the studies in the Heliyon review did not report engagement as an outcome. Some share of the null results across that literature is probably measurement design rather than product failure, though the review does not separate the two. A platform that never instrumented something cannot tell anyone whether it moved.
Two constraints change the architecture rather than the paperwork.
GDPR Article 22(1) gives a person "the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning him or her or similarly significantly affects him or her." Where a platform gates a certification or a credential that affects someone's job, that threshold is live, and Article 22(3) requires safeguards including "the right to obtain human intervention on the part of the controller, to express his or her point of view and to contest the decision." In architecture terms, the decision layer has to log why it made each routing call, expose that reason, and support a human override. Retrofitting all three into a service that only emits a score is expensive. Whether your specific decisions cross the threshold is a question for your counsel.
Explainability also sells. An L&D buyer's procurement committee asks how the system decides, and "this learner has not shown mastery of two prerequisite items for this module, here they are" answers that in a way a probability cannot. A decision layer that emits a human-readable reason alongside its choice serves the compliance requirement and the sales conversation with one mechanism.
Each phase is named by what it lets you prove.
At the end of this phase you can reconstruct any learner's session from events, and answer "where do learners leave this course" without a data pull. No model exists yet.
At the end of this phase the content team has a working authoring process, items carry competency tags, and the system routes a learner past material they have already demonstrated. This is the control condition for everything that follows.
At the end of this phase a model predicts next-item correctness better than the rules baseline on held-out data, the decision layer consumes it, and a rule-based fallback answers when inference times out.
At the end of this phase an experiment against the holdout cohort shows a difference in completion or time to competency that you would defend in front of a board.
Cost bands, team sizes and calibration timelines appear confidently in most guides on this subject. They are absent here because none of the ones we could find trace back to a source. Size each phase from four numbers you already have: items that exist versus items that need authoring, daily active learners, events per learner-hour, and competencies the graph will hold. Those drive storage, inference volume and content workload, and your own engineers' estimates against them will beat anyone's published average.
Four capabilities decide whether adaptive personalization can be built in-house, and most platform teams hold two of them.
The first is streaming data engineering: event schema design, an LRS or its equivalent, and parity between the features a model sees online and the ones it trained on offline. The second is psychometrics and learner modelling, meaning item calibration and the ability to evaluate a knowledge tracing model without the leakage the pyKT authors documented. The third is learning platform interoperability, covering xAPI or Caliper, LTI for tool integration, SCORM for legacy content, and the LMS integrations corporate buyers ask about. The fourth is experimentation design, which is the randomization and guardrail work above.
Our read is that platform engineering teams usually have interoperability covered and psychometrics not at all. That second capability is also the one hardest to hire for on a single project and easiest to bring in for a defined period, since calibration and model evaluation are bounded pieces of work with a clear end. Interoperability runs the other way, since it never stops being needed, which makes it the poorest candidate for outsourcing.
Arbisoft has built and maintained learning platforms since 2013 in the education industry as well as it contributes upstream to Open edX, Django, FastAPI and Moodle.
Enough to calibrate items, which is a smaller requirement than enough to train a deep model. A rules-based path needs no historical data at all, only tagged content and a mastery threshold. Item calibration under item response theory needs a planned sample per item, best estimated by simulating your own test design rather than copying a rule of thumb. The deep knowledge tracing models in the literature were trained on hundreds of thousands to millions of responses.
Not for the outer loop. A language model gives step-level feedback and hints well, which is inner-loop work, and the Harvard trial shows what that is worth when prompts and solutions are expert-authored. It holds no calibrated estimate of what a learner knows across a competency graph and no auditable record of how a routing decision was made. Pairing a language model for the inner loop with a knowledge model for the outer loop is the design that follows from that split.
Personalization is the broad category and adaptivity is one mechanism inside it. Personalized learning covers anything tailored to an individual, including self-selected paths, role-based curricula and preference settings. An adaptive learning platform changes the path based on evidence of what a learner has and has not mastered, which requires a learner model and assessment data feeding it continuously.
A custom schema works until data has to mean something in someone else's system. xAPI and Caliper exist so a learning record from one tool carries the same meaning in another, which matters to corporate buyers who want data in their own warehouse or LMS. Open edX converts its native tracking logs to xAPI for that reason. Starting custom with a documented transformation to xAPI is a defensible middle path.
Longer than one release, and the answer depends on cohort length. Completion cannot move faster than the time a cohort takes to finish a course, so a 12-week program yields one clean read per quarter. Leading indicators move sooner: time to first success, the drop-off point inside a module, and remediation loops per competency. Instrument those in phase one so the wait produces information.
The highest-value first move is instrumentation. Take one course, define the event schema for every interaction in it, and run both loops manually for a month, with a rules-based next-item decision and a rules-based hint, each logged with its reason. That month produces the item performance data, the calibration sample and the baseline that every later model has to beat, at a fraction of what a modelling project costs. If the rules baseline alone moves drop-off in that one course, the business case for the rest is already written.
Trusted by top platforms for our transformative solutions and exceptional results:






