# Diagnosis before generation | A 36-page SwaVid whitepaper on PAL | SwaVid

SwaVid’s detailed research paper on diagnosis-first personalised adaptive learning, durable understanding, the evidence behind PAL, and the limits of our first-party data.

Canonical: https://swavid.com/research/why-diagnosis-first

Source: https://swavid.com/research/why-diagnosis-first

# Diagnosis before generation.

## Evidence, interpretation, and our view stay separate.

## The unit of personalisation is a decision, not a description.

## Assisted performance is not the same thing as learning.

## An answer engine and a tutor solve different problems.

## One-to-one tutoring set the benchmark, not the blueprint.

## The 2-sigma result is a challenge to explain, not a claim to inherit.

## Progress should mean secured knowledge, not completed content.

## Knowledge behaves more like a graph than a playlist.

## Learning debt is local, dynamic, and repairable.

## The same wrong answer can come from different causes.

## A learner model is a history of evidence, not a personality profile.

## A tutor should know when its map is not good enough.

## PAL is a closed learning loop, not a content recommendation feed.

## Not all personalisation is equally educational.

## PAL can improve outcomes through five linked mechanisms.

## A well-designed PAL programme produced substantial gains in Delhi.

## PAL effects can survive scale, but implementation remains part of the treatment.

## Intelligent tutors tend to help, but the average hides important conditions.

## Unguided generative help can improve practice while weakening later performance.

## The right explanation depends on what the learner already knows.

## Examples teach structure only when learners process the reasoning.

## Retrieval and spacing turn a successful lesson into accessible knowledge.

## Overall understanding appears when knowledge survives a change of form.

## The learner should gradually become better at diagnosing themselves.

## Map, decide, teach, check, update.

## A live product record, not a polished research dataset.

## The evidence came in waves, not a laboratory curve.

## What learners told us before and after the profile.

## The profile looked like the child teachers knew.

## The cohort revealed a pattern worth investigating.

## Here is the line between what we know and what we do not.

## The next study must measure what remains after the lesson.

## Personalisation should increase support without increasing surveillance.

## The map is the product. Generation is how it speaks.

## Primary research and authoritative records cited.

Why personalised adaptive learning must map the learner before it generates the lesson, and why learning must be measured after the help disappears.

Published research tells us which mechanisms are plausible. Our first-party data describes current product use and profile review. SwaVid’s design principles are stated as our position. None is allowed to borrow certainty from another category.

PAL in this paper means personalised adaptive learning. It is not a public product mode or a client-selected teaching style. In SwaVid, learning strategy is selected internally from curriculum, learner evidence, session goals, and safety rules.

A system becomes a tutor only when evidence about a learner changes what happens next.

Most products called AI tutors personalise the surface. They remember a name, change the tone, generate another example, or shorten an explanation. These features can improve comfort, but they do not establish that the system has selected the right idea, difficulty, representation, or moment to teach.

SwaVid takes a stricter view. Personalisation exists only when the learner model changes an instructional decision. The decision might be to repair a prerequisite, switch from a worked example to guided practice, delay a hint, revisit a concept after a spacing interval, or stop teaching and ask for independent retrieval. If the same lesson would have been served without the learner evidence, the experience was customised, not adaptive.

This distinction is architectural. The generative model cannot be the authority on the learner, the curriculum, or the next objective. It can express an approved teaching move, but the move must be chosen by a system that can inspect evidence, represent uncertainty, enforce curriculum boundaries, and check whether learning survived without assistance.

Our test is simple: remove the learner model. If the experience stays materially the same, it is not personalised adaptive learning.

Sources: [ 1 ] , [ 3 ] , [ 5 ]

The learner must be able to retrieve, explain, and transfer an idea after the help disappears.

AI creates a measurement problem because it can raise the quality of work produced during assistance without improving the knowledge that remains afterward. A learner can copy a method, follow a sequence of hints, or accept a polished explanation and appear successful. The visible answer is better, but the underlying capability may be unchanged.

We therefore reserve the word learning for a durable change in independent capability. Immediate correctness is useful process evidence, but it is weak outcome evidence. A stronger check asks the learner to solve a fresh problem without the tutor. A stronger one still checks after a delay, changes the surface form, or asks the learner to explain why the method works.

This is why SwaVid treats support and assessment as different states. A teaching interaction may be generous with scaffolds. A mastery check must sharply limit them. Without that separation, the system evaluates the quality of its own help and mistakes that result for the learner’s understanding.

An AI tutor should be judged by the quality of the learner’s later unaided work, not by the elegance of the conversation.

Sources: [ 11 ] , [ 14 ] , [ 15 ]

The answer engine optimises the response. The tutor optimises the learner’s next state.

Optimises the answer in front of the learner.

Optimises the learner’s next state.

A learner asks about quadratic equations. An answer engine interprets the prompt and produces the clearest available explanation. A tutor first asks a different set of questions. Does the learner understand variables, signed arithmetic, factorisation, and the meaning of equality? Is the difficulty conceptual, procedural, linguistic, attentional, or simply a momentary slip?

Those questions alter the target. The correct next move may not mention quadratic equations at all. It may repair negative-number operations, use an area model to rebuild factorisation, or ask one discriminating question before teaching. The immediate query still matters, but it is evidence inside a longer model rather than an instruction the system must obey literally.

The difference becomes more important as models become more fluent. Persuasive explanations make a target error feel like success. A coherent response can conceal that the learner has been given a lesson outside their current zone of productive challenge.

The best response to the stated question can be the wrong educational response to the learner who asked it.

Sources: [ 1 ] , [ 4 ] , [ 5 ]

Bloom’s 2-sigma result describes a goal for education systems, not a guaranteed effect size for software.

Bloom’s clue was responsiveness, not the label “tutor.”

Benjamin Bloom’s 1984 paper reported that students receiving one-to-one tutoring with mastery conditions performed about two standard deviations above a conventional class. The average tutored learner was above roughly 98 percent of the control group. The result became a durable challenge: find group methods that approach the responsiveness of an excellent tutor.

The result is frequently used as marketing shorthand for AI. That leap is not warranted. Bloom did not test language models, and software does not inherit the effect of human tutoring by adopting the label tutor. The important clue is the process: continuous diagnosis, corrective feedback, appropriate practice, and enough time for mastery.

For SwaVid, 2-sigma is a design direction. It says that responsiveness matters enormously. It does not permit us to claim a 2-sigma product result without a study designed to measure one.

The honest ambition is to mechanise the responsive work of tutoring, then measure the result rather than assume it.

Sources: [ 1 ] , [ 2 ]

Bloom compared conventional instruction, mastery learning, and one-to-one tutoring. The familiar separation motivates responsive instruction. It does not establish that an AI system produces the same result.

Conventional group instruction

the average classroom

Mastery learning in groups

≈ 1 sigma above

One-to-one tutoring

≈ 2 sigma above

Interpretation: a tutor continuously aligns instruction to learner evidence. Product implication: reproduce and test the mechanism. Evidence boundary: do not advertise a 2-sigma effect without a direct, credible outcome study.

Coverage moves the timetable. Mastery changes what a learner can reliably do.

Conventional pacing treats time as fixed and learning as variable. The class advances when the calendar says to advance, even when some learners have not secured the prerequisite idea. Mastery learning reverses the relationship: the outcome is held more stable while time, explanation, practice, and corrective support are allowed to vary.

This does not require endless repetition or a punitive pass gate. It requires a defensible threshold, a targeted corrective loop, and multiple forms of evidence. A learner might show conceptual understanding before procedural fluency, or execute a procedure while being unable to explain it. The system needs to know which claim it is making.

A diagnosis-first architecture therefore represents mastery at the concept level, not the chapter-completion level. It records what evidence supports the estimate, how recent that evidence is, and what would cause the estimate to change.

The map should never say mastered merely because the learner reached the end of a lesson.

Sources: [ 2 ] , [ 17 ]

When a prerequisite is missing, every downstream success becomes harder to interpret.

School curricula are presented as sequences of chapters, but learning depends on a denser structure. Ratio supports percentage. Place value supports decimal operations. Proportional reasoning supports slope. A learner may encounter the same foundational idea across subjects and grades under different labels.

This structure explains a common classroom puzzle: a learner can follow today’s example yet fail tomorrow’s variation. The visible topic was not necessarily the point of failure. A missing prerequisite forced the learner to hold too many intermediate steps in working memory, imitate a surface pattern, or rely on a teacher cue that was absent later.

A useful learning map therefore links curriculum concepts to their likely prerequisites while remaining open to evidence. It is not enough to encode a static dependency graph. The system must ask which edge is relevant for this learner, in this problem, now.

A good map reduces detours: repair enough foundation to restore progress, then return to the learner’s goal.

Sources: [ 4 ] , [ 12 ]

A gap belongs to a concept relationship, not to a child’s identity.

Deficit labels travel too easily from one result to a broad judgment about a learner. SwaVid uses learning debt differently. It is the set of prerequisite claims that appear insufficient for a current goal, weighted by evidence and consequence. It can shrink quickly when the right concept is repaired.

The word debt is useful only if it creates action. A gap with no connection to the learner’s next objective may be low priority. A small weakness that blocks many downstream ideas may deserve immediate attention. The system should consider leverage, confidence, recency, curriculum importance, and the cost of a false diagnosis.

This makes the map directional. It does not ask only, What is weak? It asks, What is blocking progress, what evidence supports that belief, and what is the least disruptive repair?

The point of diagnosis is not to produce a more detailed report. It is to shorten the path back to meaningful learning.

Sources: [ 2 ] , [ 4 ]

Error correction without cause identification often produces another fragile rule.

Suppose two learners write the same incorrect fraction answer. One may not understand common denominators. Another may understand the concept but make an arithmetic slip. A third may misread the language. Treating the answer as a single error class leads to a generic explanation that is too basic for one learner and misses the actual problem for another.

A diagnosis-first system selects questions for information value, not merely difficulty. One carefully chosen contrast can separate a misconception from missing recall. A request to explain a step can distinguish conceptual knowledge from pattern imitation. Response time, hint use, revisions, and consistency across representations can add evidence, but none is decisive alone.

Diagnosis should remain parsimonious. The objective is not to infer a psychological biography from a short interaction. It is to distinguish the small set of hypotheses that would lead to different teaching decisions.

Useful diagnosis is decision-focused: infer only what is necessary to choose a better next action.

Sources: [ 3 ] , [ 5 ]

Knowledge tracing estimates changing mastery from imperfect observations.

Corbett and Anderson’s Bayesian Knowledge Tracing formalised a durable idea: a hidden knowledge state can be updated from observable performance while accounting for guessing and slipping. Later approaches changed the mathematics, but the core problem remains. The system never directly sees understanding. It sees answers, explanations, choices, timing, hint use, and later retention.

That distinction matters. A learner model should store claims with provenance and uncertainty. “Likely understands equivalent fractions based on two independent items” is a better state than “visual learner” or “weak at maths.” The former can be tested and revised. The latter is broad, sticky, and difficult to falsify.

SwaVid’s conceptual standard is an evidence ledger: each important estimate should point back to observations, decay when evidence becomes stale, and remain open to contradiction.

Personalisation becomes safer when the system remembers why it believes something, not only what it believes.

Sources: [ 3 ]

False certainty is more damaging than a short diagnostic pause.

Learner evidence is noisy. A correct response may be guessed. An incorrect response may reflect distraction. A learner may know a procedure in one representation and fail to recognise it in another. The system should not collapse this ambiguity into a clean but unsupported label.

Uncertainty changes policy. When confidence is low and the cost of being wrong is high, the next activity should gather information. When confidence is low but the consequence is small, the system can choose a reversible teaching move and observe the result. When a learner repeatedly contradicts the model, the model should yield.

This is also a product-design principle. The student does not need to see probabilistic machinery. They need a calm next step. Teachers and internal reviewers need enough evidence to understand why that step was chosen.

The interface can be simple because the evidence policy behind it is disciplined.

Sources: [ 3 ] , [ 5 ]

Personalised adaptive learning repeatedly measures, decides, teaches, checks, and updates.

PAL is often used loosely to describe any digital experience that differs by learner. We use the term in a narrower sense. A PAL system maintains an evidence-based learner state, selects an instructional objective and strategy, delivers an activity, measures the learner’s response, and changes the state and policy as a result.

Recommendation is one component. Personalisation might select a relevant example, but adaptation must also decide whether an example is the right instructional form. A learner who has not yet formed a concept may need a concrete representation. A learner with partial knowledge may need a completion problem. A learner who appears fluent may need delayed retrieval or transfer.

PAL should not be presented to students as a public mode they must configure. The learner’s job is to learn. Strategy selection belongs on the server, bounded by curriculum, evidence, pedagogy, safety, and session goals.

A platform is not PAL because it generates unique content. It is PAL when evidence closes the loop.

Sources: [ 6 ] , [ 7 ] , [ 8 ]

The deeper the adaptation, the more directly it can change learning.

At the shallowest level, a system changes cosmetic details: name, theme, tone, or an interest reference. These can improve engagement. The next level changes presentation: language, pace, representation, reading load, or example context. Deeper adaptation changes pedagogy: worked example, analogy, guided questioning, retrieval practice, or independent problem solving.

The deepest level changes curriculum path and timing. It identifies a prerequisite, repairs it, returns to the goal, schedules a later check, and updates the map. This level is also the riskiest because a wrong inference can send a learner away from the right material. It therefore requires stronger evidence and clearer constraints.

SwaVid’s view is additive. Surface relevance is useful, but it should sit on top of instructional adaptation rather than substitute for it.

The strongest PAL changes the learning decision while preserving the learner’s agency and curriculum destination.

Sources: [ 5 ] , [ 12 ] , [ 13 ]

The outcome is produced by the loop, not by personalisation as a vague property.

First, targeting reduces mismatch between the lesson and the learner’s current knowledge. Second, scaffolding manages the amount and type of support so the learner can make progress without outsourcing the thinking. Third, feedback arrives while the relevant reasoning is active and addresses the cause of an error. Fourth, mastery checks prevent unsupported progression. Fifth, spacing and retrieval strengthen access after the immediate lesson.

These mechanisms reinforce one another. Better targeting lowers unnecessary cognitive load. Lower load makes it easier to notice structure. A clean mastery check improves the learner model. A better model schedules more useful practice. The benefit is cumulative when the loop remains coherent across sessions.

The chain can also break. Poor diagnosis targets the wrong concept. Excessive hints create dependence. Immediate checks inflate mastery. Recommendation without follow-up loses the evidence. Any PAL claim should specify which mechanism is expected to operate and how it will be measured.

A causal mechanism is more useful than the label personalised because it tells us what to build, observe, and improve.

Sources: [ 8 ] , [ 14 ] , [ 15 ]

A 2019 randomised evaluation showed that technology can raise learning when instruction is matched to actual level.

Muralidharan, Singh, and Ganimian evaluated a technology-aided after-school programme in urban India. The programme combined computer-adaptive instruction with instructor support. In a sample of 619 students, lottery-based access produced gains of about 0.37 standard deviations in mathematics and 0.23 in Hindi after roughly 4.5 months.

The mechanism is central. Students in the sample were often several grade levels behind the curriculum, while the software personalised instruction to their learning level. The result does not prove that any adaptive product will work. It supports the proposition that matching instruction to actual knowledge can unlock learning when grade-level instruction is misaligned.

The study also evaluated a blended implementation with time, place, staff, and incentives. A chat interface deployed under different conditions should not borrow its effect size.

The right lesson for SwaVid is to preserve the level-matching mechanism and test our own delivery conditions.

Sources: [ 6 ]

A 2025 working paper reports meaningful gains after the Delhi model expanded to a much larger sample.

Muralidharan and Singh report results from an at-scale evaluation more than twenty times larger than the original study. After 18 months, the programme increased mathematics scores by about 0.22 standard deviations and Hindi scores by about 0.20. The authors estimate a 50 to 66 percent increase in learning productivity relative to the comparison condition.

The estimated effects are smaller than in the first efficacy trial but still educationally meaningful. That pattern is familiar: systems meet more heterogeneous learners, staff, schedules, and infrastructure at scale. Reliability, attendance, device access, implementation support, and local fit become part of the educational technology.

We treat these results as promising external evidence, not a SwaVid outcome. The paper is a working paper as of this review, and its programme design, operating context, and population should remain visible whenever the result is cited.

A PAL engine should be designed for real attendance, intermittent devices, multilingual classrooms, and imperfect implementation.

Sources: [ 7 ]

Meta-analyses support intelligent tutoring while warning against category-level certainty.

Ma and colleagues synthesised 107 effect sizes involving 14,321 participants. Intelligent tutoring systems outperformed teacher-led large-group instruction, non-intelligent computer instruction, and textbooks or workbooks in the analysed studies. They did not show a significant advantage over individualised human tutoring or small-group instruction.

Kulik and Fletcher reviewed 50 controlled evaluations and reported a median effect of 0.66 standard deviations over conventional instruction. The distribution was not uniform. Results depend on the subject, comparison, assessment alignment, duration, implementation, and quality of the tutor.

VanLehn’s review adds a mechanism-oriented interpretation: interaction granularity matters. Systems that engage with intermediate reasoning steps can approach more responsive forms of tutoring better than systems that wait for a final answer.

The literature justifies building serious adaptive tutors. It does not remove the obligation to evaluate each tutor.

Sources: [ 5 ] , [ 8 ] , [ 9 ]

More capable assistance can reduce the cognitive work that causes learning.

Help stays inside a path that returns thinking to the learner.

Bastani and colleagues studied generative AI use with nearly one thousand high-school mathematics students. Access to a general GPT interface improved performance during practice, but those students performed worse than controls when the assistance was removed. A more constrained GPT tutor improved supported practice much more strongly and largely mitigated the later harm, though it did not establish a positive unaided exam effect.

The result illustrates a central risk. If the model supplies the crucial step too early, the learner can produce correct work without constructing the knowledge needed to reproduce it. The tutoring experience feels successful because the visible product improves immediately.

SwaVid’s response is not to remove help. It is to make help contingent and pedagogical: ask before telling, diagnose before explaining, reveal less before more, and follow assistance with an independent check.

The tutor’s job is not to make every moment easy. It is to make difficult thinking possible and then let the learner own it.

Sources: [ 11 ]

Instruction that helps a novice can become redundant or distracting for an expert.

Cognitive load theory distinguishes the complexity inherent in a task from load created by the way instruction is presented. A novice may need a worked example because too many interacting elements would otherwise overwhelm working memory. A more knowledgeable learner may benefit from solving the same problem with less guidance.

The expertise reversal effect shows why fixed pedagogy fails. Detailed support can become redundant as prior knowledge increases. Redundant explanations split attention and consume capacity that could be used for practice, abstraction, or transfer.

PAL can operationalise this insight by adapting guidance to evidence. It can begin with a modelled solution, move to a completion problem, then remove steps and prompts. The learner model determines the fade, while independent responses determine whether the fade should continue.

Personalisation improves understanding when it changes the amount of support, not only the wording of support.

Sources: [ 12 ] , [ 13 ]

A worked solution becomes educational when the learner explains why each step is valid.

Worked examples can reduce unproductive search for novices and make a solution structure visible. But reading a correct solution is not sufficient. Learners differ in the quality of the self-explanations they generate, and those explanations predict what they learn from the same material.

An adaptive tutor can prompt the right kind of processing. It can hide a step and ask the learner to complete it, contrast two solutions, ask why a rule applies, or request a prediction before revealing the next move. These prompts are more valuable when diagnosis has identified the concept that needs attention.

Examples should also vary carefully. Surface variation helps the learner separate the deep relation from the story context. Too much variation before the core structure is stable can add noise. The sequence should respond to evidence of abstraction, not to a fixed item count.

The goal is not for the learner to remember the example. It is to recognise and reconstruct the principle elsewhere.

Sources: [ 16 ] , [ 18 ]

What feels fluent now can be inaccessible later unless the learner practises bringing it back.

Restudying often produces a strong feeling of fluency because the material is present. Retrieval practice removes that support and requires the learner to reconstruct the idea. Roediger and Karpicke showed that repeated testing can produce better delayed retention than repeated study even when study looks stronger on an immediate test.

Spacing adds a timing decision. Cepeda and colleagues reviewed 839 assessments across 317 experiments and found that the advantage of distributed practice depends jointly on the spacing interval and the desired retention interval. There is no single optimal gap for every goal.

PAL can use this relationship directly. The learner model estimates whether retrieval is likely to be effortful but possible, schedules a check, observes the result, and adjusts. A failure can trigger a short repair; success can lengthen the interval or increase transfer distance.

A learning system should remember what the learner is in danger of forgetting.

Sources: [ 14 ] , [ 15 ]

A concept is not secure if it works only in the format in which it was taught.

Near transfer changes surface details while preserving much of the structure. Farther transfer requires the learner to notice a relation across different contexts, representations, or tasks. Both are stronger evidence than repeating the original item.

A diagnosis-first system can plan transfer deliberately. It can teach a ratio with a visual model, check it numerically, then ask the learner to recognise the same relation in speed, recipes, or scale drawings. If performance collapses when the representation changes, the system has learned something specific: the first success may have been tied to a procedure or cue.

Overall understanding is therefore multidimensional. It includes accurate knowledge, connections to prerequisites, flexible representation, explanation, retrieval after delay, and transfer to a new problem. A single score can summarise progress, but it should not erase these dimensions internally.

PAL should build a connected concept that can travel, not a collection of correct responses tied to one screen.

Sources: [ 5 ] , [ 16 ]

The most valuable learner model is one the learner can increasingly participate in.

Adaptive systems can create dependence if every next step is invisibly chosen. A stronger design teaches learners to notice what they know, identify a breakdown, select a strategy, and evaluate whether the strategy worked. The system’s diagnosis becomes a scaffold for self-regulation.

This does not mean exposing every internal score. It means asking calibrated questions: How sure are you? Which step changed the answer? What prerequisite would make this easier? Would you like an example, a hint, or another attempt? The learner’s choice is itself evidence, but it also develops agency.

Over time, the system should need to do less. A learner who can accurately judge confidence, retrieve an appropriate strategy, and recover from error has gained more than topic mastery. They have gained a portable learning capability.

The endpoint of personalisation is not permanent system control. It is greater learner control.

Sources: [ 16 ]

The learner sees one calm path. The server carries the curriculum boundaries, learner evidence, strategy selection, and mastery policy. Generation is one bounded capability inside the loop.

Locate the curriculum target, prerequisites, recent evidence, and uncertainty.

Select the next objective, strategy, scaffold, and representation on the server.

Generate or retrieve a bounded activity that expresses the selected strategy.

Collect independent evidence with assistance reduced or removed.

Revise mastery, schedule retrieval, repair a gap, or advance the path.

Aggregate database review on 24 July 2026 . These counts describe use of the cognitive assessment and report workflow. They are not unique-learner counts and they are not learning-outcome measures.

Privacy boundary: this paper publishes aggregates only. It includes no names, contact details, raw answers, student-level rows, or private scoring parameters.

Monthly volume shows actual classroom and product activity from February to July 2026. The June peak demonstrates throughput under a larger burst. It should not be interpreted as a trend in learning.

Our interpretation: the workflow has been used at meaningful volume and handled a concentrated deployment period. The next evidence step is not more activity alone. It is cleaner cohort identity, matched pre and post measures, and delayed independent checks.

Reported experience remained strongly positive. Pre-report feedback averaged 4.47 across 191 responses. Post-report feedback averaged 4.66 across 32 . The distributions are useful product signals, not a matched causal comparison.

Before the report

avg 4.47 · n = 191

After the report

avg 4.66 · n = 32

What we can say: respondents were strongly positive at both moments. What we cannot say: the report caused a 0.19-point improvement, or that the rating predicts learning.

The validation team reports classroom review with 150+ students and 85 % teacher-confirmed profile agreement. This is evidence about perceived profile fit, not academic efficacy.

Share of profile reviews where the teacher said the report matched the child they see in class.

Evidence status: useful face-validity and classroom-relevance signal. The result is not independently reproducible from the aggregate product tables used for the submission snapshot, and it does not measure grades, retention, or transfer.

Nine aggregate trait scores describe the assessed cohort. Structured thinking is lowest in this snapshot. That observation can shape questions and product hypotheses, but it does not establish a prerequisite deficit or a stable trait in any learner.

Amber marks the cohort&#x27;s weakest trait: structured, step-by-step thinking

Product implication: use the pattern to choose discriminating follow-up evidence and to improve explanations. Do not convert a cohort average into an individual label or a claim about fixed learning style.

Strong research communication makes the denominator, comparison, and outcome visible. This is the current boundary as of 24 July 2026 .

Our commitment: future product growth will not silently upgrade a feasibility signal into an efficacy claim. Each new claim needs a measure designed for that claim.

A credible evaluation programme should tell us not only whether scores moved, but why, for whom, under what conditions, and whether the knowledge remained available.

Preferred design: preregister outcomes, use an active comparison, keep assessors blind where practical, report attrition, and publish null or mixed results alongside positive ones.

A learner model is powerful because it shapes opportunity. That makes restraint, access control, and reversibility part of the pedagogy.

Collect the minimum evidence required for an educational decision. Prefer observable learning behaviour over speculative personality inference. Separate private scoring parameters and policy from the student client. Limit access by role, retain an audit trail for consequential decisions, and make deletion and correction operationally possible.

Adaptation must not become a ceiling. A low-confidence estimate should not permanently route a learner into easier work. The system should offer routes to disconfirm the model, periodically probe readiness, and surface persistent conflicts for human review.

Generative output remains bounded by curriculum context, age appropriateness, session scope, and a server-owned strategy. The product can explain why it chose a next step without exposing private internals or pretending that the estimate is certain.

Personalisation should remain a revisable hypothesis in service of the learner, never a permanent verdict about them.

Personalised adaptive learning can improve outcomes when it reduces instructional mismatch, protects productive thinking, verifies independent mastery, and returns at the right time. Research supports those mechanisms and shows promising results in specific implementations. It also shows why unconstrained assistance can create the appearance of learning without the durable capability.

SwaVid’s unique position is therefore diagnosis before generation. The learner model chooses the target. The curriculum bounds the path. The strategy selects the teaching move. Generation expresses it. Independent evidence decides whether to repair, revisit, or advance.

Our current first-party evidence shows real use, positive reported experience, and promising profile agreement. It does not yet prove learning gains. That is not a footnote. It is the next research task.

Links resolve to the publisher, DOI record, ERIC, NBER, or another primary bibliographic source. First-party figures are identified on their exhibit pages and reviewed as of 24 July 2026 .

Suggested citation: SwaVid. (2026). Diagnosis before generation: Why personalised adaptive learning must map the learner before it generates the lesson. SwaVid Research.

> “An answer engine responds to the question in front of the learner. A tutor responds to the learner in front of the question.”

- [ 1 ] Bloom, B. S. (1984). The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring. Educational Researcher, 13(6), 4–16.
- [ 2 ] Bloom, B. S. (1968). Learning for Mastery. Evaluation Comment, 1(2).
- [ 3 ] Corbett, A. T., & Anderson, J. R. (1995). Knowledge Tracing: Modeling the Acquisition of Procedural Knowledge. User Modeling and User-Adapted Interaction, 4, 253–278.
- [ 4 ] Ausubel, D. P. (1968). Educational Psychology: A Cognitive View.
- [ 5 ] VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist, 46(4), 197–221.
- [ 6 ] Muralidharan, K., Singh, A., & Ganimian, A. J. (2019). Disrupting Education? Experimental Evidence on Technology-Aided Instruction in India. American Economic Review, 109(4), 1426–1460.
- [ 7 ] Muralidharan, K., & Singh, A. (2025). Improving Schooling Productivity through Computer-Aided Personalization: Experimental Evidence from Rajasthan. NBER Working Paper 34205.
- [ 8 ] Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. (2014). Intelligent Tutoring Systems and Learning Outcomes: A Meta-Analysis. Journal of Educational Psychology, 106(4), 901–918.
- [ 9 ] Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research, 86(1), 42–78.
- [ 10 ] Banerjee, A. V. et al. (2017). From Proof of Concept to Scalable Policies: Challenges and Solutions, with an Application. Journal of Economic Perspectives, 31(4), 73–102.
- [ 11 ] Bastani, H. et al. (2025). Generative AI Can Harm Learning. Proceedings of the National Academy of Sciences, 122(17).
- [ 12 ] Sweller, J. (1988). Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science, 12(2), 257–285.
- [ 13 ] Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The Expertise Reversal Effect. Educational Psychologist, 38(1), 23–31.
- [ 14 ] Roediger, H. L., & Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science, 17(3), 249–255.
- [ 15 ] Cepeda, N. J. et al. (2006). Distributed Practice in Verbal Recall Tasks: A Review and Quantitative Synthesis. Psychological Bulletin, 132(3), 354–380.
- [ 16 ] Chi, M. T. H. et al. (1989). Self-Explanations: How Students Study and Use Examples in Learning to Solve Problems. Cognitive Science, 13(2), 145–182.
- [ 17 ] Kulik, C. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of Mastery Learning Programs: A Meta-Analysis. Review of Educational Research, 60(2), 265–299.
- [ 18 ] Renkl, A., Stark, R., Gruber, H., & Mandl, H. (1998). Learning from Worked-Out Examples: The Effects of Example Variability and Elicited Self-Explanations. Contemporary Educational Psychology, 23(1), 90–108.

## Key Links

- [[ 1 ]](https://doi.org/10.3102/0013189X013006004)
- [[ 3 ]](https://doi.org/10.1007/BF01099821)
- [[ 5 ]](https://doi.org/10.1080/00461520.2011.611369)
- [[ 11 ]](https://doi.org/10.1073/pnas.2422633122)
- [[ 14 ]](https://doi.org/10.1111/j.1467-9280.2006.01693.x)
- [[ 15 ]](https://doi.org/10.1037/0033-2909.132.3.354)
- [[ 1 ]](https://doi.org/10.3102/0013189X013006004)
- [[ 4 ]](https://openlibrary.org/books/OL26705958M/Educational_Psychology_A_Cognitive_View)
- [[ 5 ]](https://doi.org/10.1080/00461520.2011.611369)
- [[ 1 ]](https://doi.org/10.3102/0013189X013006004)
- [[ 2 ]](https://eric.ed.gov/?id=ED053419)
- [[ 2 ]](https://eric.ed.gov/?id=ED053419)
- [[ 17 ]](https://doi.org/10.3102/00346543060002265)
- [[ 4 ]](https://openlibrary.org/books/OL26705958M/Educational_Psychology_A_Cognitive_View)
- [[ 12 ]](https://doi.org/10.1207/s15516709cog1202_4)
- [[ 2 ]](https://eric.ed.gov/?id=ED053419)
- [[ 4 ]](https://openlibrary.org/books/OL26705958M/Educational_Psychology_A_Cognitive_View)
- [[ 3 ]](https://doi.org/10.1007/BF01099821)
- [[ 5 ]](https://doi.org/10.1080/00461520.2011.611369)
- [[ 3 ]](https://doi.org/10.1007/BF01099821)
- [[ 3 ]](https://doi.org/10.1007/BF01099821)
- [[ 5 ]](https://doi.org/10.1080/00461520.2011.611369)
- [[ 6 ]](https://doi.org/10.1257/aer.20171112)
- [[ 7 ]](https://www.nber.org/papers/w34205)
- [[ 8 ]](https://doi.org/10.1037/a0037123)
- [[ 5 ]](https://doi.org/10.1080/00461520.2011.611369)
- [[ 12 ]](https://doi.org/10.1207/s15516709cog1202_4)
- [[ 13 ]](https://doi.org/10.1207/S15326985EP3801_4)
- [[ 8 ]](https://doi.org/10.1037/a0037123)
- [[ 14 ]](https://doi.org/10.1111/j.1467-9280.2006.01693.x)
- [[ 15 ]](https://doi.org/10.1037/0033-2909.132.3.354)
- [[ 6 ]](https://doi.org/10.1257/aer.20171112)
- [[ 7 ]](https://www.nber.org/papers/w34205)
- [[ 5 ]](https://doi.org/10.1080/00461520.2011.611369)
- [[ 8 ]](https://doi.org/10.1037/a0037123)
- [[ 9 ]](https://doi.org/10.3102/0034654315581420)
- [[ 11 ]](https://doi.org/10.1073/pnas.2422633122)
- [[ 12 ]](https://doi.org/10.1207/s15516709cog1202_4)
- [[ 13 ]](https://doi.org/10.1207/S15326985EP3801_4)
- [[ 16 ]](https://doi.org/10.1207/s15516709cog1302_1)
- [[ 18 ]](https://doi.org/10.1006/ceps.1997.0959)
- [[ 14 ]](https://doi.org/10.1111/j.1467-9280.2006.01693.x)
- [[ 15 ]](https://doi.org/10.1037/0033-2909.132.3.354)
- [[ 5 ]](https://doi.org/10.1080/00461520.2011.611369)
- [[ 16 ]](https://doi.org/10.1207/s15516709cog1302_1)
- [[ 16 ]](https://doi.org/10.1207/s15516709cog1302_1)
- [See the diagnosis-first learning map](https://swavid.com/learning-debt-identifier)
- [Read the first-party evidence case study](https://swavid.com/case-studies/apple-global-school)
- [Bloom, B. S. (1984). The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring. Educational Researcher, 13(6), 4–16.](https://doi.org/10.3102/0013189X013006004)
- [Bloom, B. S. (1968). Learning for Mastery. Evaluation Comment, 1(2).](https://eric.ed.gov/?id=ED053419)
- [Corbett, A. T., & Anderson, J. R. (1995). Knowledge Tracing: Modeling the Acquisition of Procedural Knowledge. User Modeling and User-Adapted Interaction, 4, 253–278.](https://doi.org/10.1007/BF01099821)
- [Ausubel, D. P. (1968). Educational Psychology: A Cognitive View.](https://openlibrary.org/books/OL26705958M/Educational_Psychology_A_Cognitive_View)
- [VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist, 46(4), 197–221.](https://doi.org/10.1080/00461520.2011.611369)
- [Muralidharan, K., Singh, A., & Ganimian, A. J. (2019). Disrupting Education? Experimental Evidence on Technology-Aided Instruction in India. American Economic Review, 109(4), 1426–1460.](https://doi.org/10.1257/aer.20171112)
- [Muralidharan, K., & Singh, A. (2025). Improving Schooling Productivity through Computer-Aided Personalization: Experimental Evidence from Rajasthan. NBER Working Paper 34205.](https://www.nber.org/papers/w34205)
- [Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. (2014). Intelligent Tutoring Systems and Learning Outcomes: A Meta-Analysis. Journal of Educational Psychology, 106(4), 901–918.](https://doi.org/10.1037/a0037123)
- [Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research, 86(1), 42–78.](https://doi.org/10.3102/0034654315581420)
- [Banerjee, A. V. et al. (2017). From Proof of Concept to Scalable Policies: Challenges and Solutions, with an Application. Journal of Economic Perspectives, 31(4), 73–102.](https://www.nber.org/papers/w22746)
- [Bastani, H. et al. (2025). Generative AI Can Harm Learning. Proceedings of the National Academy of Sciences, 122(17).](https://doi.org/10.1073/pnas.2422633122)
- [Sweller, J. (1988). Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science, 12(2), 257–285.](https://doi.org/10.1207/s15516709cog1202_4)
- [Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The Expertise Reversal Effect. Educational Psychologist, 38(1), 23–31.](https://doi.org/10.1207/S15326985EP3801_4)
- [Roediger, H. L., & Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science, 17(3), 249–255.](https://doi.org/10.1111/j.1467-9280.2006.01693.x)
- [Cepeda, N. J. et al. (2006). Distributed Practice in Verbal Recall Tasks: A Review and Quantitative Synthesis. Psychological Bulletin, 132(3), 354–380.](https://doi.org/10.1037/0033-2909.132.3.354)
- [Chi, M. T. H. et al. (1989). Self-Explanations: How Students Study and Use Examples in Learning to Solve Problems. Cognitive Science, 13(2), 145–182.](https://doi.org/10.1207/s15516709cog1302_1)
- [Kulik, C. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of Mastery Learning Programs: A Meta-Analysis. Review of Educational Research, 60(2), 265–299.](https://doi.org/10.3102/00346543060002265)
- [Renkl, A., Stark, R., Gruber, H., & Mandl, H. (1998). Learning from Worked-Out Examples: The Effects of Example Variability and Elicited Self-Explanations. Contemporary Educational Psychology, 23(1), 90–108.](https://doi.org/10.1006/ceps.1997.0959)