by Kenneth Sung, Boon Lishi Lisa, Hoe Sion Hee, Shawn Neo Zhe Ming
Section 1: Context – The GAINS Initiative and the Questions It Raised
Generative Artificial Intelligence (GenAI) is rapidly reshaping how feedback is conceived, generated, and delivered in classrooms worldwide. It offers significant possibilities to help students receive and understand timely feedback and expectations and mitigate demands on teacher workload and limited curriculum time. On the other hand, it also presents challenges about the integrity of assessment, accuracy of feedback and the alignment to intended curriculum outcomes. It is within this context that the Global AI Nexus of Schools (GAINS) was established in 2025. Founded by Hwa Chong Institution, GAINS establishes various Network Learning Communities (NLCs) to bring together a growing community of schools working together to study how GenAI can be meaningfully and responsibly integrated into teaching, learning, and assessment practices.
Naval Base Secondary School (NBSS) is glad to be part of the AI-Enhanced Assessment Feedback NLC, working with seven schools in Singapore and Indonesia to better understand different approaches in the deployment of GenAI in assessment feedback practices across subject disciplines. It provides a collaborative platform through which teachers across participating schools can critically examine GenAI-enhanced approaches in the feedback cycle through projects carried out by participating schools.
The Chemistry team from NBSS embarked on a term-long study on the use of GenAI in teaching and learning with Secondary 3 students, with the details and findings detailed in the later part of this article.
As the team embarked on this journey, our teachers also share concerns about the use of GenAI in teaching and learning. These concerns fell broadly into three categories.
Accuracy of feedback and alignment to syllabus outcomes
GenAI tools are trained on vast, general datasets capable of providing detailed feedback and guidance, but may not reliably reflect the specific requirements and expectations of a given subject or level. This increases the risk of feedback information that is technically accurate but may be misaligned to curriculum outcomes or students’ readiness. In addition, the possibility of AI hallucinations leading to inaccurate feedback information misguiding students meant teachers will still have to be in the “loop” to check for such inaccuracies.
Gaps in pedagogical know-how on effective integration of Gen AI
The effective use of GenAI requires the know-how in designing meaningful learning experiences for the students, and for students to be familiar with the interaction with AI tools at their disposal. Teachers were still grappling with some foundational design questions, such as when should GenAI feedback support or supplement students’ learning instead of teachers’ comments, as well as designing experiences that encourages students to think critically about the AI feedback rather than passive acceptance. The use of GenAI feedback, which can exist on multiple online platforms, to feed forward in students’ learning also requires teachers to ensure feedback information are intentionally captured for students’ reference.
Scalability and sustainability of integrating AI in the feedback cycle
Technical know-how for the adequate setting up of GenAI tools – be it on the Student Learning Space (SLS) or third-party platforms – remain an area teachers can be more familiar with. The ability to accurately set up instructions and parameters for the GenAI tools, and teaching students to write effective prompts meant time is needed to build and guide the use of the GenAI tools. This brings up the question of whether the use of GenAI can be scaled across different classes, levels and subject teams in the department and school, and whether this can be done without significantly adding on to the workload of the teachers.
These concerns were not set aside — they were brought directly into the design of the department’s inquiry, examined through the TEASA framework (Fig. 1) introduced by A/P Kelvin Tan to the participating schools in the NLC (see table below). The framework, summarised below, provided a structured lens through which to evaluate whether AI-enhanced feedback was genuinely serving students and teachers.
Fig. 1 TEASA Framework
Section 2: The Foundation — Four Years of Building a Feedback Ecosystem
It is within this landscape of possibility and uncertainty that the Chemistry Team of Science Department at Naval Base Secondary School began its GAINS journey. But to understand what the department brought to that journey — and why its approach to AI-enhanced feedback looked the way it did — it is necessary to go back four years, to a time before AI was part of the conversation at all.
Good feedback has always been the goal. The question the Science Department kept asking was: how do we make it more visible, more timely, and more useful — for both teachers and students?
The answer unfolded over four years, through a progressive transformation of the department’s pedagogical and feedback ecosystem — moving deliberately from static, paper-based tracking to an agile, AI-enhanced feedback loop. This evolution was neither rushed nor technology-driven for its own sake. It began deliberately in 2023 with student artefacts and a simplified RIF survey to establish a reliable baseline of student readiness, before introducing Interactive Feedback Cover Sheets in 2024 to build dialogic feedback habits between teachers and students. Each phase laid the groundwork for the next, ensuring that both teachers and students were genuinely ready for what followed.
By 2025, with students and teachers now fluent in the language of feedback, the department was ready to scale its impact through technology. A suite of AI tools — SAFA, SALis, Data Assistant, and Mizou — were introduced alongside Structured Learning Logs, each serving a distinct pedagogical purpose: providing immediate rubric-aligned feedback, clustering misconceptions, guiding dialogic problem-solving, and helping students consolidate their learning. By 2026, the focus shifted to stacking these tools strategically, with students trained in prompt engineering and teachers leveraging real-time data for deeper conceptual guidance. Throughout all four phases, comprehensive data triangulation across RIF surveys, Weighted Assessment scores, and Focus Group Discussions ensured that decisions were always grounded in evidence rather than intuition.
Seen against the backdrop of the TEASA framework, the department’s four-year journey reads almost as an unwitting preparation for the questions GAINS was designed to ask. The RIF survey had been building the evidence base needed to evaluate trustworthiness and equity long before those terms entered the conversation. The Interactive Feedback Cover Sheets had been cultivating the dialogic habits that human agency requires. The progressive introduction of AI tools had been testing the boundaries of scope, scale, and sustainability in real classrooms with real students. When the opportunity arose to join GAINS and form a Network Learning Community, the chemistry team was not starting from scratch — it was bringing four years of principled inquiry to bear on a set of questions it had, in many ways, already been living.
The NLC became the platform through which its accumulated knowledge of AfL theory, feedback literacy, and data-informed practice could be tested and scaled in the context of AI-enhanced learning. Rather than asking “how do we use AI in our classrooms,” the department was able to ask a far more sophisticated question: “how does AI deepen and extend the feedback ecosystem we have already built?” It is this question that shaped the department’s GAINS journey.
Section 3: The Experiment
The GAINS initiative presented the Science Department with a timely and significant opportunity — not to start afresh, but to build purposefully on the strong pedagogical foundation it had spent four years constructing. By the time GAINS was introduced, the department was well-positioned to engage with it meaningfully. Teachers had already developed a shared assessment language grounded in the Triangulated Model of AfL and the Four Boxes of Feedback Literacy. PLT units were functioning with growing autonomy, guided by a culture of evidence-based inquiry. The RIF survey had given the department a reliable instrument for tracking student dispositions toward feedback over time. And years of structured professional development — both school-wide and department-led — had built the kind of professional trust and collaborative rigour that is rarely assembled quickly.
The department was candid with itself that it was still finding its footing — experimenting with different combinations of tools, observing how students responded, and iterating based on what the data revealed. Not every approach worked as anticipated, and not every tool proved equally effective in every context. This willingness to experiment without the pressure of having all the answers was itself a product of the safe learning culture the department had spent years building — and it was precisely this spirit of principled experimentation that laid the groundwork for the more sophisticated multi-tool integration that would follow in 2026.
The Physics PLT team piloted their AI-enhanced feedback approach in 2025 across Secondary 3 Express and 3NA Science (Physics) classes. The lesson was designed around a transfer of learning task integrating concepts across Kinematics, Forces, and Dynamics, where students first attempted a question individually before refining their answers through iterative dialogue with Mizou, an AI chatbot, ahead of a teacher-led consolidation. Mizou’s ability to provide immediate, personalised feedback regardless of student readiness level meant that every student could engage meaningfully with the task at their own pace — something that would have been difficult to replicate manually across a full class (Fig. 2a). As students arrived at the consolidation phase with more developed responses, teachers were freed from addressing basic misconceptions and could direct their attention toward deeper conceptual guidance. Beyond the academic gains, the team also observed a subtler but equally significant benefit: Mizou created a low-stakes environment where students felt safe to make mistakes and think aloud without fear of judgement — a quality that is difficult to engineer in a traditional classroom setting, and one that speaks directly to the department’s longer-term goal of building genuine feedback literacy.
They encountered real challenges along the way. Crafting effective prompts required many rounds of refinement before Mizou behaved consistently with the team’s pedagogical intentions. Student responses were also mixed — those accustomed to receiving direct answers found the bot’s Socratic approach frustrating, while students with weaker initial responses found the volume of feedback overwhelming rather than actionable. These reactions were revealing in themselves, offering an honest measure of where students still were in their readiness to engage productively with feedback. Looking ahead, with the introduction of the Learning Assistant feature in SLS (SALiS), the team now plans to consolidate their approach within a single integrated platform, streamlining the experience while preserving the real-time, personalised feedback that made the Mizou pilot worthwhile.
Fig. 2a Physics PLT leveraged the AI chatbot Mizou to provide immediate, non-judgmental feedback.
The Biology PLT team piloted their approach in 2025 with Secondary 3E Science (Biology) classes on the topic of Digestion and Enzymes. The impetus was a familiar classroom observation: students consistently produced short answers riddled with misconceptions or missing key scientific terminology. The team turned to SAFA and AFA within the SLS platform to provide immediate, customised feedback on individual responses, while the Data Assistant summarised common errors across the class — significantly reducing the time teachers spent manually reviewing submissions (Figure 2b). Most students were able to understand the improvements needed, and the combination of tools allowed for more targeted differentiated instruction than had previously been possible.
The Biology team’s experience surfaced important limitations that underscore why teacher oversight remains indispensable. AI-generated feedback was only as reliable as the suggested answer prompts and commands teachers had carefully constructed — and even then, the tools occasionally awarded full credit to answers that missed critical keywords or contained outright misconceptions. These were not minor edge cases but recurring issues that required teachers to review all submissions rather than defer entirely to the AI’s judgement. Far from being a drawback, the team reframed these moments as valuable teaching opportunities — using instances of AI inaccuracy to help students develop a more critical and discerning relationship with the technology itself.
Fig. 2b Biology PLT utilized SLS’s SAFA, AFA, and Data Assistant to address student misconceptions.
By 2025, all Science PLT units had pioneered the integration of AI-enhanced feedback pedagogy (Fig. 2) — a collective milestone that reflected the department’s growing confidence in navigating this emerging landscape. The pilots had surfaced important early lessons: that student readiness could not be assumed, that tool trustworthiness required active management, and that the teacher’s role was not diminished by AI but made more purposeful. These lessons set the stage for the Chemistry PLT’s more rigorous and extended inquiry as part of the GAINS NLC project — a new chapter in the department’s AI-enhanced feedback journey.
Section 4: The GAINS NLC Project – Methodology
The Chemistry PLT team began designing and implementing their project in 2025 before continuing to refine their approach in 2026 as part of the GAINS NLC. Their project involved 93 students across three Secondary 3 Science (Chemistry) classes, structured around two inquiry cycles anchored in the feedback cycle of Feed Up, Feedback, and Feed Forward. Teachers began each cycle by establishing clear success criteria to set learning goals — the Feed Up stage. Students then attempted an SLS quiz and used SAFA and AFA to identify their learning gaps, documenting their difficulties in a Learning Log before engaging with the Sidekick AI chatbot to address those gaps. Performance data from quizzes and formal assessments allowed both teachers and students to track progress over time — the Feed Forward stage.
The two cycles were deliberately sequenced to build student agency progressively. In Cycle 1 on the Structure and Properties of Materials, students used the COSTAR framework and Bloom’s Taxonomy to scaffold their prompts for Sidekick (Figure 3a) — providing the structure needed for students still developing their confidence with AI-assisted learning. By Cycle 2 on Acids and Bases, this scaffolding was gradually released as students transitioned toward more natural, dialogic prompting.
Fig. 3a Artifact of a student using COSTAR framework
As illustrated in Figure 3b below, each cycle captured three interconnected moments in a student’s learning journey: AFA provided immediate, annotated feedback on the student’s response; Sidekick guided the student through their misconceptions dialogically without prematurely giving away answers; and the student consolidated what they had learned through their own self-evaluation in the learning log — embodying the department’s vision of a truly self-regulated feedback cycle. To evaluate impact, the team drew on a rich body of evidence — student artefacts, AI interaction logs, assessment performance, pre and post RIF surveys, and multi-ability focus group discussions — reflecting the same rigorous, triangulated approach to data that had characterised the department’s inquiry from the very beginning.
Feedback from AFA and student’s record of learning difficulties captured in learning log
Insights from Sidekick after interaction with student
Fig. 3b Artefacts of the AI-enhanced lesson, showing AFA feedback, student learning log entries, Sidekick interaction insights, and student self-evaluation.
Data Revealed — Findings
What Students Felt About AI Feedback
Students completed a pre-survey on their receptivity to teacher feedback and a post-survey on their receptivity to AI feedback. The findings were striking (Fig. 4). Across all students, there was a significant decrease in ratings when shifting from teacher to AI feedback, with the sharpest drop in Instrumental Attitudes (from 3.86 to 3.25) and an overall engagement drop from 3.80 to 3.40. Students consistently rated AI feedback lower than personalised teacher feedback, suggesting that while AI provides rapid responses, students still prioritise the perceived value and emotional connection of human-led instruction.
Fig. 4 Data from Pre and Post RIF surveys
A much clearer and more detailed picture emerged when the data was analysed by readiness group, together with the Focus Group Discussion findings. High Readiness students showed significant decreases across almost all domains and were the most critical of AI’s instrumental value — they felt the feedback was sometimes too generic, and expressed frustration at “dead ends” where the chatbot could not clarify next steps. Some reported distrust due to instances of AI hallucination and expressed a clear preference for a blended approach with teacher feedback. Moderate Readiness students showed significant drops in the Experiential, Cognitive, and Instrumental domains, though their Behavioural Engagement remained stable. Interestingly, many appreciated the emotional safety of interacting with the chatbot — as one student noted, the AI “has no feelings,” which made them feel less judged during learning struggles. Low Readiness students were the least affected group overall, finding AI feedback nearly as acceptable as teacher feedback, and viewing Sidekick as a kind of personal tutor for quick access — though they still sought teacher assurance for deeper understanding.
Relating Performance data to Perception data
High Readiness students were generally able to articulate their learning gaps clearly in the chatbot, note down key learning points and self-evaluate against the learning objectives. Their performance data reflected this — in Cycle 2, HR students saw their scores jump from 63.4% on the initial SLS Quiz to 77.3% on the Feed Forward Quiz taken immediately after AI-assisted feedback, and continued improving to 82.5% in FA2 and 82.7% in WA2. The AI reinforced an already solid conceptual foundation and supported self-regulated learning.
Moderate Readiness (MR) students found SAFA and AFA useful in identifying their mistakes, and the chatbot was effective at triggering engagement. However, they were frequently frustrated by circular questioning — instances where Sidekick explained concepts back to them rather than providing clear next steps. Despite this, their scores showed a steady upward trend in Cycle 2, from 54.9% to 70.0% to 82.8%, though there was a slight dip to 66.3% by WA2, suggesting that AI-triggered engagement did not always translate into long-term retention without further teacher intervention.
Low Readiness (LR) students presented the most complex picture. While SAFA and AFA helped them highlight missing keywords and break problems into manageable steps, many tended to use Sidekick as a quick fix rather than engaging in deep processing. Significant misconceptions persisted across the quiz, FA2 and WA2. One student, for instance, continued to commit the same error regarding alkali reactions despite having recorded this mistake in his Learning Log. Their Cycle 2 scores reflected this pattern — moving from 36.6% to 55.6% to 59.1%, before dipping back to 44.3% in WA2.
Fig. 12 Performance Data and Its Relation to RIF findings
How Students Talked to the AI
Beyond performance data, we also analysed how students’ interaction patterns with Sidekick evolved across the two cycles. In Cycle 1, approximately 48% of students used the COSTAR framework, which acted as a kind of training wheels — forcing students to provide context and state their objectives clearly. This resulted in high-quality, comprehensive initial responses from the chatbot that established strong conceptual foundations, with an average dialogue depth of 3.56 turns, mostly focused on seeking clarification of definitions.
By Cycle 2, students had shifted to direct, task-oriented queries — troubleshooting specific mistakes in chemical formulas and engaging in sustained back-and-forth exchanges to work through multi-step problems. The average dialogue depth increased slightly to 3.65 turns, but the nature of the interaction had changed meaningfully: students had, in effect, learned how to talk to the AI. The chatbot evolved from functioning like a textbook — providing definitions — to functioning more like a coaching partner, helping students troubleshoot reactions in real time. This progression from structured scaffold prompting to natural dialogic prompting was one of the most encouraging developments we observed across the two cycles.
Section 6: Making Sense of it All – The TEASA Lens
The Chemistry PLT’s GAINS NLC project was not simply an experiment in deploying new tools — it was a deliberate attempt to address the very concerns teachers had raised about AI-enhanced feedback, and to do so in a way that made feedback more agentic, more impactful, and more sustainable. Mapping the project against the TEASA framework helps illuminate how each design decision responded to a specific concern, and what the findings reveal about the conditions under which AI-enhanced feedback genuinely improves learning and the learner.
Trustworthiness
One of the most pressing concerns teachers raised was whether AI-generated feedback could be trusted — not just in terms of factual accuracy, but across the affective, behavioural, cognitive, and utility dimensions of learning. This concern proved well-founded. The Chemistry team found that SAFA and AFA, while effective at identifying surface-level errors, occasionally awarded full credit to answers containing misconceptions or missing critical keywords. High Readiness students reported distrust stemming from instances of AI hallucination, and expressed frustration when the chatbot reached dead ends without offering clear next steps.
The Learning Log played a crucial role in addressing this concern. Rather than leaving students to rely solely on AI-generated responses, the Learning Log required students to actively record their difficulties, document what they had learned from Sidekick, and self-evaluate against the learning objectives. This meant that even when AI feedback was imperfect, students were not passive recipients — they were prompted to interrogate and consolidate what the AI had told them, creating a layer of critical engagement that mitigated the risk of misinformation going unexamined. Teachers remained in the loop throughout, reviewing submissions and using interaction logs to identify where AI judgement had fallen short.
Going forward, the department proposes formalising this vigilance through an Evaluate-Simulate-Refine protocol — treating AI tool configuration with the same rigour applied to assessment design. Before any tool reaches the classroom, teachers would first evaluate whether the system prompt is pedagogically sound and aligned to learning objectives, then simulate student interactions by stress-testing the chatbot with the kinds of incomplete or misconception-laden responses real students are likely to produce, and finally refine the instructions iteratively until outputs are consistently reliable. This cycle is not a one-time exercise but an ongoing discipline, recognising that as topics, student cohorts, and AI capabilities change, the process of calibration must continue.
Efficiency, Effectiveness, and Equity
The concern about scalability — whether AI-enhanced feedback could reduce teacher workload without compromising quality — was directly addressed by the project’s design. Tools like SAFA and AFA provided immediate, annotated feedback on individual student responses at a scale no single teacher could replicate manually, while the Data Assistant surfaced common errors across the class, significantly reducing the time teachers spent reviewing submissions. This freed teachers to focus their attention on deeper conceptual guidance rather than routine error correction — a meaningful shift in how teacher effort was deployed.
Effectiveness, however, proved more nuanced. The data showed that AI’s impact varied significantly by student readiness, with High Readiness students sustaining strong gains, Moderate Readiness students showing improvement that required teacher reinforcement to consolidate, and Low Readiness students struggling to translate AI-assisted engagement into lasting learning. On equity, the findings were sobering: AI feedback is not a leveller in itself. Without deliberate teacher scaffolding, the risk is that AI-enhanced feedback inadvertently widens rather than narrows the gap between students. The Learning Log was one mechanism designed to address this — by requiring all students, regardless of readiness, to document their learning and self-evaluate, it created a structured opportunity for deeper processing that the AI interaction alone could not guarantee. Even so, the data showed that Low Readiness students continued to commit the same errors despite having recorded them in their logs, underscoring that the log is a necessary but not sufficient condition for equity — teacher intervention remains indispensable.
Looking ahead, improving accessibility is a priority that equity demands. Current tools rely primarily on text-based input, which can disadvantage students who struggle to articulate their thinking in writing. Exploring interfaces that support photo or audio inputs would lower the barrier to engagement and ensure that AI-assisted feedback is genuinely inclusive across different learner profiles. For Moderate and Low Readiness students in particular, the department recognises that AI-triggered engagement does not automatically translate into long-term retention — and that teacher intervention must be deliberately planned rather than left to chance.
Ambition
The concern about whether AI feedback would remain task-limited — correcting immediate errors without building enduring understanding or genuine learner independence — was central to the project’s design ambitions. The two-cycle structure was deliberately sequenced to move students beyond surface correction toward self-regulation. In Cycle 1, the COSTAR framework and Bloom’s Taxonomy scaffolded students’ prompts, ensuring they provided context and stated their objectives clearly. By Cycle 2, this scaffolding was gradually released as students transitioned toward more natural, dialogic prompting — a progression that was one of the most encouraging developments the team observed.
The Learning Log was the primary vehicle for translating AI feedback into enduring learning. By requiring students to consolidate what they had learned through their own self-evaluation — rather than simply receiving and moving on — the log embodied the department’s vision of a truly self-regulated feedback cycle. The data bore this out most clearly for High Readiness students, who used Sidekick not merely to correct errors but to monitor their own understanding and drive their learning forward. For lower readiness students, the ambition of building genuine independence remained a work in progress.
One of the clearest lessons from the pilots was that students need repeated, sustained exposure to AI-assisted feedback across multiple topics before they can engage with it fluently and independently. Single-lesson pilots, however well-designed, are insufficient to consolidate the habits of mind that effective AI interaction requires. Extending the opportunities for students to interact with chatbots across topics and over time — building longer runways for productive AI dialogue habits to develop — will be critical to ensuring that the ambition of genuine learner independence is not just a design intention but a measurable outcome.
Scope, Scale, and Sustainability
The concern about whether AI-enhanced feedback could be sustained across classes, levels, and subject teams without significantly adding to teacher workload was tested directly by the project. Involving 93 students across three Secondary 3 Chemistry classes, the project demonstrated that the approach was viable at a meaningful scale within a single subject team. The use of SLS-integrated tools — SAFA, AFA, and Sidekick — meant that the technical infrastructure was already familiar to both teachers and students, reducing the setup burden that had been identified as a barrier to scalability.
That said, the project was candid about the real costs involved. Crafting effective prompts for Sidekick required many rounds of refinement, and the proposed Evaluate-Simulate-Refine protocol, while essential for trustworthiness, represents a genuine investment of teacher time. The Learning Log also required careful design to ensure it served its pedagogical purpose rather than becoming a mechanical compliance exercise. Sustainability, the team concluded, depends not on minimising teacher effort but on ensuring that effort is purposefully directed — and that the tools are configured well enough to justify the investment. The department’s next step is to extend this approach beyond Chemistry to the Physics and Biology PLTs, using the lessons from the GAINS project to build a shared departmental resource for AI-enhanced feedback that can be adapted across subject teams without each team having to start from scratch.
Agency (Human)
Perhaps the most significant dimension of the TEASA framework — and the one that most directly addresses the concern about passive acceptance of AI feedback — is human agency. The project’s defining insight, which the team termed the “Teacher Effect,” was that AI cannot on its own convert feedback interactions into sustained learning growth. It is the teacher who bridges the gap between a student receiving feedback and a student genuinely acting on it. This was borne out clearly in the FA2 and WA2 results, where the most significant gains were observed not simply where AI tools were used, but where teachers had intervened deliberately and responsively based on what the data revealed.
This played out in telling ways across the PLT teams. In the Physics classes, when students grew frustrated by Mizou’s refusal to provide direct answers, it was the teacher who stepped in to reframe the struggle — validating the discomfort, explaining the pedagogical intent behind the chatbot’s Socratic approach, and helping students recognise that the difficulty they were experiencing was itself a sign of learning. In the Chemistry classes, some students began passively copying Sidekick’s responses without internalising them, and it was the teacher who interrupted that passivity with strategic questions that restored the reflective intent of the task.
A recurring challenge across all three pilots was students’ tendency to approach AI feedback as a faster route to the correct answer rather than a genuine opportunity for reflection. Addressing this requires more than reminders — it requires a sustained cultural shift in how students understand the purpose of feedback itself, and it places the teacher at the centre of that shift. The Learning Log was designed precisely to counter passive acceptance, requiring students to articulate their own understanding in their own words rather than transcribe the AI’s output. For teachers, the real-time interaction logs surfaced which students were struggling and where thinking had broken down, enabling them to move through the classroom with purpose rather than intuition. Going forward, the department is deliberate in ensuring that as AI tools become more capable and more embedded in classroom practice, the teacher’s role is not gradually marginalised in favour of efficiency. AI should function as the initial scaffolding layer — handling immediate, individualised feedback at scale — while teachers remain the essential providers of human assurance, contextual judgement, and the kind of relational support that no algorithm can replicate.
Section 8: Conclusion
Four years ago, the Science Department made a quiet but consequential decision: that feedback was not a finishing touch to be added after learning, but the very substance through which learning happens. That conviction shaped every subsequent choice — the structured protocols, the gradual introduction of AI tools, and the rigorous inquiry that followed.
Three years of pilots and one rigorous NLC project have brought the department to a further clarity: the question is not whether AI belongs in the feedback loop, but how thoughtfully that loop is designed. Student readiness must be cultivated, tool trustworthiness must be actively built, and the teacher’s judgement must remain at the centre — not despite the presence of AI, but because of what AI makes newly visible.
It is this conviction that will guide the department’s work in the years ahead.
