The Kirkpatrick Model is a four-level framework for evaluating whether training actually worked. Level 1 (Reaction) asks whether learners found the training relevant and engaging. Level 2 (Learning) asks whether knowledge or skills actually increased. Level 3 (Behavior) asks whether people apply what they learned on the job, 30, 60, or 90 days later. Level 4 (Results) asks whether an organizational outcome moved because of the training. Donald Kirkpatrick introduced the model in 1959. In 2026, Vanessa Milara Alzate formally expanded it to include the performance environment as an explicit component and extended its application beyond L&D to enterprise performance intelligence. According to ATD’s 2025 research, the Kirkpatrick Model is used by 71% of organizations surveyed, making it the most common training evaluation framework in use. The reason most implementations fail: organizations measure Level 1 (the post-training survey) and stop there, never reaching the Level 3 and Level 4 data that shows whether training created any real change.
Key Highlights of Kirkpatrick Model of Training Evaluation ATD’s 2025 State of the Industry research found the Kirkpatrick Model is used by 71% of organizations it is the most commonly used training evaluation framework globally Kirkpatrick Partners confirmed that in 2026, Vanessa Milara Alzate formally expanded the model to include the performance environment and extend application beyond L&D to full enterprise performance intelligence ATD’s 2016 research (still the most widely cited in the field) found only 35% of organizations evaluate at Level 3 (Behavior) and fewer than 10% evaluate at Level 4 (Results) the two most valuable levels The New World Kirkpatrick Model, developed by Dr. Jim Kirkpatrick, reverses the design process: start with Level 4 (what business result do we need?) and work backward to Level 1 (what experience will support that learning?) Konstantly’s 2026 Kirkpatrick guide recommends measuring Level 3 starting within 2 to 4 weeks of training, not at 90 days, because by 90 days the behavior window has often closed Level 3 “Required Drivers,” the reinforcement systems (coaching, feedback loops, accountability mechanisms) that must surround training to produce lasting behavior change, are as important as the training content itself The Kirkpatrick Model of training evaluation answers the question every L&D leader eventually faces: did this training actually change anything? Not whether people enjoyed it. Not whether they passed the quiz. Whether, weeks after the program ended, they do their job differently as a result.
That question sounds obvious, but most organizations never answer it. They collect post-training surveys (Level 1), file the results, and move on. According to research from Devlin Peck’s comprehensive 2026 Kirkpatrick guide , ATD found that only 35% of organizations evaluate at Level 3 (Behavior) and fewer than 10% evaluate at Level 4 (Results). They are measuring the least useful data and skipping the most useful.
This guide covers all four levels with worked examples, the 2026 updates to the framework including Vanessa Milara Alzate’s expansion, the New World Kirkpatrick Model’s backward planning approach, the most common measurement tools for each level, and the limitations every L&D professional should know before designing an evaluation plan. For organizations applying this framework to Generative AI training programs specifically, NextAgile’s Gen AI Training Services incorporate Level 3 and Level 4 measurement into the program design itself, not as an afterthought.
Background: Who Created the Kirkpatrick Model and When Donald Kirkpatrick, an American professor and president of the American Society for Training and Development (now ATD), developed the four-level framework through a series of articles published in 1959. He later formalized it in his 1994 book “Evaluating Training Programs.”
The model was updated significantly by his son, Dr. Jim Kirkpatrick, who developed the “New World Kirkpatrick Model” in the 2010s. The New World version added the concept of “Required Drivers” (the organizational conditions that make behavior transfer possible) and reversed the design process to start with Level 4 outcomes first.
In 2026, Kirkpatrick Partners confirmed that Vanessa Milara Alzate expanded the model further to include the performance environment as an explicit component and moved the framework’s utility beyond learning and development to apply to the entire enterprise, supporting what she called “enterprise performance intelligence.”
The model’s greatest strength, as Kirkpatrick Partners notes, “is its simplicity and flexibility, making it adaptable across industries, organizations, and program types.” Across 65 years of evolution, the four-level vocabulary has remained consistent.
The Four Levels of the Kirkpatrick Model Level 1: Reaction What it measures: Did participants find the training engaging, relevant, and favorably delivered?
When to measure: During and immediately after training.
Why it matters: Level 1 is not just a satisfaction score. A training that learners find irrelevant will produce poor behavior transfer regardless of content quality. If learners disengage because the delivery is poor or the examples do not match their context, Level 2 and 3 outcomes suffer. Level 1 data is early warning, not success confirmation.
The common mistake: Treating Level 1 as an evaluation of the trainer rather than of the learning experience. The Kirkpatrick Model, as Ardent Learning’s guide explains, “encourages survey questions that concentrate on the learner’s takeaways” rather than on the facilitator’s performance.
What to measure at Level 1:
Perceived relevance to current job (critical: if learners do not see immediate applicability, transfer fails) Confidence in applying what was covered Specific ideas they plan to try Rating of pacing, content clarity, and examples Sample Level 1 questions:
“How relevant was this training to your current role and daily work? (1-5)” “What is one thing you plan to do differently in the next 30 days as a result of today?” “What would have made this program more useful for your specific situation?” The third question is the most valuable and least commonly used: it surfaces the gap between what was designed and what learners actually needed.
A real example: A financial services organization running a Generative AI literacy program for senior managers found their Level 1 data revealing: managers rated content quality 4.2/5 but relevance to their specific team context only 2.8/5. That mismatch predicted exactly the behavior transfer failure they observed at Level 3. Designing for role-specific examples would have been the correct intervention, and Level 1 caught it early enough to adjust the program before the second cohort. NextAgile’s Generative AI Workshop for Enterprise is designed specifically to prevent this role-context mismatch by tailoring examples to participants’ actual workflows.
Level 2: Learning What it measures: Did participants’ knowledge, skills, attitudes, confidence, or commitment actually increase during the program?
When to measure: At the end of training and, ideally, at the start (as a baseline).
Why it matters: Level 2 confirms whether the learning experience delivered any knowledge transfer at all. You cannot change behavior (Level 3) without first changing knowledge or skill (Level 2). A program with strong Level 1 scores but flat Level 2 scores has an instructional design problem, not a facilitation problem.
The pre/post assessment design: The most rigorous Level 2 measurement uses a pre-test at the start of training and a post-test immediately after. The delta between the two scores is the attributable learning gain. Without the pre-test, you cannot distinguish between what learners already knew and what the training actually taught them.
What to measure at Level 2:
Knowledge tests (multiple choice, scenario-based) Skills demonstrations (observed role-plays, task completion) Confidence self-assessment before and after Attitude or intention shifts toward applying new approaches A real example: A manufacturing company ran safety procedure training across three sites. Level 2 post-tests showed an average score increase of 22 percentage points across all sites. But Site C showed only an 8-point increase. Investigating, the L&D team found that Site C’s shift supervisors had pre-briefed their teams on the test content before training, inflating pre-test scores and making the delta appear small. The discovery led to a redesign using scenario-based assessment that could not be coached in advance.
For training programs in agile delivery, the Agile and Scrum Masterclass measures Level 2 through scenario-based assessments where participants apply Scrum events to a simulated sprint, not just recall definitions.
Level 3: Behavior What it measures: Are participants applying what they learned on the job, and to what degree?
When to measure: Starting 2 to 4 weeks after training, with continued observation over the following months.
Why it is the most critical level: Level 3 is where training either delivers value or does not. All the Level 1 and Level 2 data in the world tells you nothing about whether the business changed. Level 3 is where you find out.
Why most organizations fail at Level 3: The measurement requires following up. Post-training surveys are easy to automate. Level 3 requires someone to observe, ask, or review performance records weeks after the program ended, when the L&D team has moved on to the next program. Konstantly’s 2026 guide recommends starting Level 3 measurement within 2 to 4 weeks, not at 90 days, because the window for reinforcement and behavior adjustment is shorter than most organizations assume.
Required Drivers: The Missing Piece Most Organizations Ignore
The New World Kirkpatrick Model introduced the concept of Required Drivers: the organizational conditions that must exist for training to transfer into behavior. Training alone rarely produces behavior change. The conditions around training do.
Required Drivers include:
Reinforcement: Managers who follow up, ask about application, and recognize new behaviors Feedback loops: Systems that let learners know whether they are applying skills correctly Accountability: Processes that hold people responsible for applying what they learned Coaching: Access to support during the application period when learners face real implementation challenges If Required Drivers are absent, even excellent training produces minimal Level 3 results. This is why organizations often blame the training content when the real problem is the absence of management reinforcement after the program ends.
What to measure at Level 3:
360-degree feedback from managers and peers Direct observation of new behaviors in meetings, deliverables, or interactions Self-assessment against a behavioral checklist with specific examples required Review of work outputs (code quality metrics, customer satisfaction scores, project delivery data) A real example: A technology company trained 80 product owners in advanced backlog management techniques. Level 1 and Level 2 scores were strong. At the 60-day Level 3 check, only 34% of participants reported regular use of the specific techniques covered. Investigation revealed the issue: line managers were unaware their direct reports had attended the training and were therefore not reinforcing the new behaviors or creating opportunities to practice them. The organization added a mandatory 30-minute manager briefing before each cohort, and Level 3 application rates for the next cohort improved to 68%.
This Required Drivers principle connects directly to how high-performing teams are built : the conditions around behavior matter as much as the behaviors themselves.
Level 4: Results What it measures: Did an organizational outcome improve because of the training?
When to measure: 3 to 6 months after training, using business metrics that were baselined before the program began.
Why it is the hardest and most valuable level: Level 4 is where training connects to business impact. It is also where attribution becomes difficult: many organizational factors affect business results simultaneously. The question is whether training was a material contributor.
The isolation challenge: How do you prove that a sales training program improved revenue rather than the new marketing campaign that launched in the same quarter? The New World Kirkpatrick Model does not require perfect causal isolation. It requires reasonable attribution through a combination of data patterns, timing, and participant-reported application evidence.
Methods for Level 4 measurement:
Control group comparison (trained vs. untrained cohort with similar characteristics) Trend line analysis (did the metric trajectory change at the point training was delivered?) Participant testimony with specific attribution (“I applied X from the training when Y happened, which led to Z outcome”) Leading indicators established at Level 3 that predict Level 4 results What to measure at Level 4 (by training type):
Training Type Example Level 4 Metrics Sales training Revenue per rep, deal close rate, pipeline conversion Leadership development Team engagement scores, retention rates, 360 feedback improvement Agile/Scrum training Sprint velocity, defect rate, release frequency GenAI training for enterprises Prompts per task reduction, AI tool adoption rate, hours saved per week Customer service training CSAT score, first-contact resolution rate, escalation rate
A real example: An insurance company ran a six-month leadership development program for 40 middle managers. Level 4 measurement at 6 months showed: team engagement scores for participants’ direct reports improved 11 percentage points versus a comparison group of similar teams, voluntary turnover in participants’ teams dropped from 18% to 11% annualized, and quality audit pass rates improved 9 points. These results justified the program continuation and secured budget for the next cohort.
NextAgile’s approach to corporate leadership training builds Level 4 measurement into program design from the start, defining the business metric that training is designed to move before the first session runs.
The New World Kirkpatrick Model: Work Backward from Level 4 The most important innovation in the New World Kirkpatrick Model is reversing the design sequence. Traditional training design moves forward: design the content, deliver the training, measure the results. New World design moves backward:
Step 1: Start at Level 4. What business result do we need to move? Define the specific, measurable organizational outcome that this training is meant to support.
Step 2: Define Level 3 behaviors. What would employees need to do differently on the job to drive that Level 4 result? These become the behavioral learning objectives.
Step 3: Identify Required Drivers. What organizational conditions (manager reinforcement, feedback systems, practice opportunities) must be in place to make those Level 3 behaviors happen?
Step 4: Design Level 2 learning. What knowledge, skill, and confidence do employees need to be able to perform those behaviors?
Step 5: Design Level 1 experience. What learning experience will build that knowledge and skill in a way that learners find relevant and engaging?
This backward design process, also called “starting with the end in mind,” ensures that every design decision links directly to a business outcome. Programs designed this way are significantly more likely to show Level 3 and Level 4 results because the learning objectives were derived from business outcomes rather than from content preferences.
For organizations designing training programs using this approach, how to develop agile training plan guide covers backward-design planning applied specifically to agile capability programs.
Kirkpatrick vs. Phillips ROI Model The Phillips ROI Model (sometimes called the Level 5 model) adds a financial ROI calculation to the Kirkpatrick framework. Level 5 asks: what is the monetary return on the training investment, after isolating the training’s contribution and accounting for program costs?
The Phillips model is more rigorous and more complex. It requires monetizing the Level 4 outcomes (converting reduced turnover or faster processing into dollar values) and applying an ROI formula. It is most appropriate for high-cost programs where formal financial justification is required.
For most organizational training programs, the four Kirkpatrick levels provide sufficient rigor to make smart investment decisions. The Phillips extension is valuable when a CFO or board specifically requires financial ROI justification for a significant L&D investment.
How to Apply the Kirkpatrick Model to GenAI Training in 2026 Applying the Kirkpatrick framework to Generative AI training programs requires thinking specifically about what Level 3 and Level 4 look like for AI skill development:
Level 1 (Reaction): Did participants find the AI content relevant to their actual role? Generic AI literacy content frequently fails this test for non-technical roles. Role-specific examples are critical.
Level 2 (Learning): Can participants demonstrate specific AI skills, not just describe them? Level 2 for GenAI training should include live demonstrations: “use a prompt engineering technique to reduce the time to complete this task by X%.”
Level 3 (Behavior): Are participants using AI tools in their daily work 30 days later? Track AI tool adoption rates and prompt quality metrics, not just self-reported usage.
Level 4 (Results): Did AI adoption reduce time on specific task categories, improve output quality scores, or enable the team to take on higher-value work? Define this metric before the program launches.
NextAgile’s Gen AI Training Services and Generative AI Foundations Workshop are designed with Kirkpatrick-aligned measurement built into the delivery framework, so organizations can report on behavior change and business impact, not just learner satisfaction.
The Main Criticisms of the Kirkpatrick Model Being balanced requires acknowledging what critics get right:
Causal links are assumed, not proven. The model implies that Level 1 leads to Level 2 leads to Level 3 leads to Level 4. Research has shown these links are weaker than the model suggests. High Level 1 scores do not reliably predict Level 2 gains, and high Level 2 scores do not guarantee Level 3 transfer.
Level 3 and 4 are rarely measured in practice. Despite the model being widely known, the ATD data is stark: most organizations never actually implement the most valuable levels. The simplicity of the model’s concept does not translate into simplicity of implementation.
The performance environment matters more than training alone. The 2026 expansion by Vanessa Milara Alzate addressed this directly: the performance environment (manager support, resource availability, organizational culture) is as important as training design. The original model understated this.
These criticisms do not undermine the model’s value. They clarify where it needs to be supplemented: with rigorous causal design where possible, and with explicit attention to the organizational conditions that make behavior transfer happen.
Conclusion The Kirkpatrick Model works when organizations commit to all four levels, not just Level 1. The model is simple. The measurement discipline it requires is not. Most L&D programs invest heavily in designing engaging Level 1 and Level 2 experiences, then measure nothing after the program ends. The result is a growing body of expensive training with no evidence of impact and no way to justify continued investment.
Three decisions that will change how your next program is designed: define the Level 4 business metric before you write a single learning objective; plan the Required Drivers (manager briefings, follow-up coaching, accountability checkpoints) as part of the program design rather than as an afterthought; and schedule Level 3 measurement for 30 days post-training, not 90. If you are designing a Generative AI training program for your organization and want Level 3 and Level 4 measurement built in from the start, NextAgile’s Gen AI Training Services and corporate leadership training programs are designed with that accountability framework embedded.
Frequently Asked Questions 1.What are the 4 levels of the Kirkpatrick Model?
The four levels are: Level 1 (Reaction) did learners find the training relevant and engaging? Level 2 (Learning) did knowledge or skills actually increase? Level 3 (Behavior) are learners applying new skills on the job 30-90 days later? Level 4 (Results) did an organizational metric move because of the training? Each level builds on the one before it and becomes progressively harder but more valuable to measure.
2.Who created the Kirkpatrick Model?
Donald L. Kirkpatrick, an American professor and president of ATD (then the American Society for Training and Development), developed the four-level framework through articles published in 1959. His son, Dr. Jim Kirkpatrick, updated it into the “New World Kirkpatrick Model” in the 2010s, adding Required Drivers and backward design principles. In 2026, Vanessa Milara Alzate expanded the framework to include the performance environment and enterprise-wide application.
3.Why do most organizations only use Level 1?
Because Level 1 (the post-training survey) is the easiest data to collect: it happens immediately after training while everyone is still in the room. Level 3 requires following up with learners and their managers weeks later. Level 4 requires baselining business metrics before the program and measuring them months afterward. Most L&D teams do not have the bandwidth, authority, or baseline data to do this consistently. Designing Level 3 and 4 measurement into the program from the start, rather than adding it afterward, is the only reliable fix.
4.What is the “New World Kirkpatrick Model”?
The New World Kirkpatrick Model reverses the design sequence: instead of designing from Level 1 forward, it starts with Level 4 (what business outcome do we need?) and works backward through Level 3 (what behaviors drive that outcome?) to Level 2 (what learning supports those behaviors?) to Level 1 (what experience builds that learning?). It also adds “Required Drivers” the organizational reinforcement systems (manager coaching, accountability mechanisms) that must exist for training to transfer into behavior.
5.How do you measure Level 3 behavior change?
Level 3 is measured through: 360-degree feedback from managers and peers, direct observation during work, self-assessment against a behavioral checklist with specific examples required (not just yes/no ratings), and review of work output metrics that would change if behavior changed. The timing matters: start Level 3 measurement 2 to 4 weeks post-training, not at 90 days. By 90 days, the opportunity to reinforce and correct early application attempts has often passed.
6.How does the Kirkpatrick Model apply to AI training programs? For AI training programs, Level 3 measures whether participants are actually using AI tools in their daily work 30 days after training, and whether the quality of their AI application (prompt quality, task completion time, output accuracy) has improved. Level 4 measures whether AI adoption reduced time on specific task categories, improved output quality scores, or enabled the team to take on higher-value work. Define both metrics before the program launches, not afterward. NextAgile’s approach to training development teams on generative AI embeds these measurement principles into the program design.
Anuj Ojha is Co-Founder & Consulting Head at NextAgile. Anuj has designed & led multiple turnkey transformation journeys across industries, domains & geographies and has 16+ years of experience as an agile practitioner. He has worked with CXOs, CTOs & Key Leaders to translate their business objectives on the ground, contextualizing org transformations and creating buy-in across level, leading a team of coaches/consultants to implement agility across 150+ teams & trained more than 12k team members. Anuj’s core area of interest is business agility & working with leaders & teams to achieve long term sustainable, Agile culture & mindset.