Behavioral Retrieval-Augmented Tutoring
Grounding a Multi-Agent LLM Math Tutor in Annotated Tutor–Student Dialogue
DOI:
https://doi.org/10.58445/rars.4229Keywords:
Intelligent tutoring systems, LLMs, multi-agent systems, RAG, dialogue tutoring, talk movesAbstract
Intelligent tutoring systems automate the adaptation of mathematics instruction for learners. This paper introduces BR-Tutor, a multi-agent mathematics tutoring architecture designed to ground a Tutor Agent in authentic pedagogical discourse rather than static textbook content. While large language models excel at mathematical problem-solving, their default disposition is to supply answers prematurely, which undermines the Socratic guided-discovery process essential to effective tutoring. To bridge this gap, the proposed system enhances an existing LangGraph-based architecture with a threshold-gated behavioral retrieval module. This module indexes 3,375 talk-move-annotated tutor-student exchanges from the Eedi QATD-2k corpus and injects the closest matching historical tutor response into the agent's prompt to model appropriate pedagogical strategies. Deployed on open-weight language models via the Groq inference platform, BR-Tutor was evaluated on a 50-problem sample from the standard MathDial benchmark. The system achieved a 94% task-success rate (Success@5) alongside a 46% answer disclosure rate (Telling@5), with both metrics rapidly saturating by the second conversational turn. Ultimately, this research highlights a persistent dissociation between raw mathematical correctness and necessary pedagogical restraint in open-weight models, emphasizing the critical role of behavioral grounding in intelligent tutoring systems.
References
Bloom, B. S.: The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher 13(6), 4–16 (1984). https://doi.org/10.3102/0013189X013006004
VanLehn, K.: The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist 46(4), 197–221 (2011). https://doi.org/10.1080/00461520.2011.611369
Wood, D., Bruner, J. S., Ross, G.: The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry 17(2), 89–100 (1976). https://doi.org/10.1111/j.1469-7610.1976.tb00381.x
Chi, M. T. H., Siler, S. A., Jeong, H., Yamauchi, T., Hausmann, R. G.: Learning from human tutoring. Cognitive Science 25(4), 471–533 (2001). https://doi.org/10.1207/s15516709cog2504_1
Kapur, M.: Productive failure. Cognition and Instruction 26(3), 379–424 (2008). https://doi.org/10.1080/07370000802212669
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., Mariman, R.: Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122(26), e2422633122 (2025). https://doi.org/10.1073/pnas.2422633122
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., et al.: Training language models to follow instructions with human feedback. In: Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pp. 27730–27744 (2022).
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., et al.: ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences 103, 102274 (2023). https://doi.org/10.1016/j.lindif.2023.102274
Macina, J., Daheim, N., Chowdhury, S. P., Sinha, T., Kapur, M., Gurevych, I., Sachan, M.: MathDial: A dialogue tutoring dataset with rich pedagogical properties grounded in math reasoning problems. In: Findings of the ACL: EMNLP 2023, pp. 5602–5621 (2023). https://doi.org/10.18653/v1/2023.findings-emnlp.372
Chudziak, J. A., Kostka, A.: AI-Powered Math Tutoring: Platform for Personalized and Adaptive Education. arXiv preprint arXiv:2507.12484 (2025). https://arxiv.org/abs/2507.12484
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., Cao, Y.: ReAct: Synergizing reasoning and acting in language models. In: Proc. 11th Int. Conf. Learning Representations (ICLR 2023) (2023).
Edge, D., Trinh, H., Cheng, N., Bradley, J., Chao, A., Mody, A., Truitt, S., Metropolitansky, D., Ness, R. O., Larson, J.: From local to global: A graph RAG approach to query-focused summarization. Microsoft Research technical report, arXiv:2404.16130 (2024). https://arxiv.org/abs/2404.16130
Zent, M., Smith, D., Woodhead, S.: PIIvot: A lightweight NLP anonymization framework for question-anchored tutoring dialogues. In: Proc. 2025 Conf. Empirical Methods in Natural Language Processing (EMNLP 2025), pp. 27479–27488. ACL, Suzhou, China (2025).
Johnson, J., Douze, M., Jégou, H.: Billion-scale similarity search with GPUs. IEEE Transactions on Big Data 7(3), 535–547 (2021). https://doi.org/10.1109/TBDATA.2019.2921572
Reimers, N., Gurevych, I.: Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In: Proc. 2019 Conf. Empirical Methods in Natural Language Processing and 9th Int. Joint Conf. Natural Language Processing (EMNLP-IJCNLP), pp. 3982–3992 (2019). https://doi.org/10.18653/v1/D19-1410
Graesser, A. C., Wiemer-Hastings, K., Wiemer-Hastings, P., Kreuz, R.: AutoTutor: A simulation of a human tutor. Cognitive Systems Research 1(1), 35–51 (1999). https://doi.org/10.1016/S1389-0417(99)00005-4
Heffernan, N., Heffernan, C.: The ASSISTments ecosystem: Building a platform that brings scientists and teachers together for minimally invasive research on human learning and teaching. International Journal of Artificial Intelligence in Education 24, 470–497 (2014). https://doi.org/10.1007/s40593-014-0024-x
van Hoeve, M., Doorman, M., Veldhuis, M.: Fostering a growth mindset in secondary mathematics classrooms in the Netherlands. Research in Mathematics Education, 1–22 (2023). https://doi.org/10.1080/14794802.2023.2241433
Sonkar, S., Liu, N., Mallick, D. B., Baraniuk, R. G.: CLASS: A design framework for building intelligent tutoring systems based on learning science principles. In: Findings of the ACL: EMNLP 2023, pp. 1941–1961 (2023). https://doi.org/10.18653/v1/2023.findings-emnlp.130
Anderson, J. R., Corbett, A. T., Koedinger, K. R., Pelletier, R.: Cognitive tutors: Lessons learned. Journal of the Learning Sciences 4(2), 167–207 (1995). https://doi.org/10.1207/s15327809jls0402_2
Koedinger, K. R., Anderson, J. R., Hadley, W. H., Mark, M. A.: Intelligent tutoring goes to school in the big city. International Journal of Artificial Intelligence in Education 8, 30–43 (1997).
Corbett, A. T., Anderson, J. R.: Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction 4(4), 253–278 (1994). https://doi.org/10.1007/BF01099821
Piech, C., Bassen, J., Huang, J., Ganguli, S., Sahami, M., Guibas, L. J., Sohl-Dickstein, J.: Deep knowledge tracing. In: Advances in Neural Information Processing Systems 28 (NeurIPS 2015) (2015).
Graesser, A. C., Person, N. K., Magliano, J. P.: Collaborative dialogue patterns in naturalistic one-to-one tutoring. Applied Cognitive Psychology 9(6), 495–522 (1995). https://doi.org/10.1002/acp.2350090604
Nye, B. D., Graesser, A. C., Hu, X.: AutoTutor and family: A review of 17 years of natural language tutoring. International Journal of Artificial Intelligence in Education 24(4), 427–469 (2014).
Kumar, P.: Large language models (LLMs): Survey, technical frameworks, and future challenges. Artificial Intelligence Review 57 (2024). https://doi.org/10.1007/s10462-024-10888-y
Wang, X., Xu, X., Zhang, Y., Hao, S., Jie, W.: Exploring the impact of artificial intelligence application in personalized learning environments: Thematic analysis of undergraduates' perceptions in China. Humanities and Social Sciences Communications 11(1) (2024). https://doi.org/10.1057/s41599-024-04168-x
Mishra, T., Sutanto, E., Rossanti, R., Pant, N., Ashraf, A., Raut, A., Uwabareze, G., Oluwatomiwa, A., Zeeshan, B.: Use of large language models as artificial intelligence tools in academic research and publishing among global clinical researchers. Scientific Reports 14(1) (2024). https://doi.org/10.1038/s41598-024-81370-6
Chen, Y., Ding, N., Zheng, H. T., Liu, Z., Sun, M., Zhou, B.: Empowering private tutoring by chaining large language models. In: Proc. 33rd ACM Int. Conf. Information and Knowledge Management (CIKM '24), pp. 354–364 (2024). https://doi.org/10.1145/3627673.3679665
Taneja, K., Maiti, P., Kakar, S., Guruprasad, P., Rao, S., Goel, A. K.: Jill Watson: A virtual teaching assistant powered by ChatGPT. In: Artificial Intelligence in Education (AIED 2024), Part I, pp. 324–337. Springer (2024).
Park, M., Kim, S., Lee, S., Kwon, S., Kim, K.: Empowering personalized learning through a conversation-based tutoring system with student modeling. In: Ext. Abstracts CHI Conf. Human Factors in Computing Systems (CHI EA '24), pp. 1–10 (2024). https://doi.org/10.1145/3613905.3651122
Liu, Z., Yin, S. X., Lee, C., Chen, N. F.: Scaffolding language learning via multi-modal tutoring systems with pedagogical instructions. In: 2024 IEEE Conf. Artificial Intelligence (CAI), pp. 1258–1265 (2024).
Tack, A., Piech, C.: The AI teacher test: Measuring the pedagogical ability of Blender and GPT-3 in educational dialogues. In: Proc. 15th Int. Conf. Educational Data Mining (EDM 2022) (2022).
Wang, R. E., Zhang, Q., Robinson, C., Loeb, S., Demszky, D.: Bridging the novice-expert gap via models of decision-making: A case study on remediating math mistakes. In: Proc. 2024 Conf. North American Chapter of the ACL (NAACL 2024) (2024).
Jiang, Y., Zhang, M., Yin, X., Jin, S., Lu, S., Ying, Z., Yu, Z., Kong, X.: EduGuardBench: A holistic benchmark for evaluating the pedagogical fidelity and adversarial safety of LLMs as simulated teachers. Proceedings of the AAAI Conference on Artificial Intelligence 40(37), 31356–31364 (2026). https://doi.org/10.1609/aaai.v40i37.40399
Kakar, S., Maiti, P., Taneja, K., Nandula, A., Nguyen, G., Zhao, A., Nandan, V., Goel, A. K.: Jill Watson: Scaling and deploying an AI conversational agent in online classrooms. In: Intelligent Tutoring Systems (ITS 2024), Part I, pp. 78–90. Springer (2024).
Shi, X., Zhang, C., Zhu, Y., Zhang, X., Luo, Y.: Beyond pedagogical principles: Multi-horizon preference optimization for efficient Socratic tutoring. In: Proc. 64th Annual Meeting of the ACL (ACL 2026), Long Papers, pp. 11289–11306. San Diego, CA, USA (2026). https://aclanthology.org/2026.acl-long.518/
Chudziak, J. A., Wawer, M.: ElliottAgents: A natural language-driven multi-agent system for stock market analysis and prediction. In: Proc. 38th Pacific Asia Conf. Language, Information and Computation (PACLIC 38), pp. 961–970. Tokyo, Japan (2024).
Chudziak, J. A., Cinkusz, K.: Towards LLM-augmented multiagent systems for agile software engineering. In: Proc. 39th IEEE/ACM Int. Conf. Automated Software Engineering (ASE '24), pp. 2476–2477 (2024).
Kostka, A., Chudziak, J. A.: Synergizing logical reasoning, long-term memory, and collaborative intelligence in multi-agent LLM systems. In: Proc. 38th Pacific Asia Conf. Language, Information and Computation (PACLIC 38). Tokyo, Japan (2024).
Viswanathan, N., Meacham, S., Adedoyin, F. F.: Enhancement of online education system by using a multi-agent approach. Computers and Education: Artificial Intelligence 3, 100057 (2022). https://doi.org/10.1016/j.caeai.2022.100057
Park, J. S., O'Brien, J., Cai, C. J., Morris, M. R., Liang, P., Bernstein, M. S.: Generative agents: Interactive simulacra of human behavior. In: Proc. 36th Annual ACM Symp. User Interface Software and Technology (UIST '23), pp. 1–22 (2023). https://doi.org/10.1145/3586183.3606763
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Zhang, C., et al.: MetaGPT: Meta programming for a multi-agent collaborative framework. In: Proc. 12th Int. Conf. Learning Representations (ICLR 2024) (2024).
Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., et al.: ChatDev: Communicative agents for software development. In: Proc. 62nd Annual Meeting of the ACL (ACL 2024), Long Papers (2024).
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., Yih, W.: Dense passage retrieval for open-domain question answering. In: Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP 2020), pp. 6769–6781 (2020). https://doi.org/10.18653/v1/2020.emnlp-main.550
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al.: Retrieval-augmented generation for knowledge-intensive NLP tasks. In: Advances in Neural Information Processing Systems 33 (NeurIPS 2020), pp. 9459–9474 (2020).
Li, X., Henriksson, A., Duneld, M., Nouri, J., Wu, Y.: Supporting teaching-to-the-curriculum by linking diagnostic tests to curriculum goals: Using textbook content as context for retrieval-augmented generation with large language models. In: Artificial Intelligence in Education (AIED 2024), pp. 118–132. Springer (2024). https://doi.org/10.1007/978-3-031-64302-6_9
Guan, X., Zhang, L. L., Liu, Y., Shang, N., Sun, Y., Zhu, Y., Yang, F., Yang, M.: rStar-Math: Small LLMs can master math reasoning with self-evolved deep thinking. In: Proc. 42nd Int. Conf. Machine Learning (ICML 2025), PMLR, vol. 267, pp. 20640–20661 (2025).
Trinh, T. H., Wu, Y., Le, Q. V., He, H., Luong, T.: Solving olympiad geometry without human demonstrations. Nature 625(7995), 476–482 (2024). https://doi.org/10.1038/s41586-023-06747-5
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., et al.: Language models are few-shot learners. In: Advances in Neural Information Processing Systems 33 (NeurIPS 2020), pp. 1877–1901 (2020).
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., Zhou, D.: Chain-of-thought prompting elicits reasoning in large language models. In: Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pp. 24824–24837 (2022).
Rubin, O., Herzig, J., Berant, J.: Learning to retrieve prompts for in-context learning. In: Proc. 2022 Conf. North American Chapter of the ACL (NAACL 2022) (2022).
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., Zettlemoyer, L.: Rethinking the role of demonstrations: What makes in-context learning work? In: Proc. 2022 Conf. Empirical Methods in Natural Language Processing (EMNLP 2022) (2022).
Macina, J., Daheim, N., Wang, L., Sinha, T., Kapur, M., Gurevych, I., Sachan, M.: Opportunities and challenges in neural dialog tutoring. In: Proc. 17th Conf. European Chapter of the ACL (EACL 2023), pp. 2357–2372 (2023). https://doi.org/10.18653/v1/2023.eacl-main.173
Suresh, A., Jacobs, J., Harty, C., Perkoff, M., Martin, J. H., Sumner, T.: The TalkMoves dataset: K-12 mathematics lesson transcripts annotated for teacher and student discursive moves. In: Proc. 13th Language Resources and Evaluation Conf. (LREC 2022), pp. 4654–4662 (2022).
Michaels, S., O'Connor, C., Resnick, L. B.: Deliberative discourse idealized and realized: Accountable talk in the classroom and in civic life. Studies in Philosophy and Education 27(4), 283–297 (2008).
Demszky, D., Liu, J., Mancenido, Z., Cohen, J., Hill, H., Jurafsky, D., Hashimoto, T.: Measuring conversational uptake: A case study on U.S. math education. In: Proc. 59th Annual Meeting of the ACL and 11th Int. Joint Conf. Natural Language Processing (ACL-IJCNLP 2021) (2021).
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., Scialom, T.: Toolformer: Language models can teach themselves to use tools. In: Advances in Neural Information Processing Systems 36 (NeurIPS 2023) (2023).
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., et al.: ToolLLM: Facilitating large language models to master 16000+ real-world APIs. In: Proc. 12th Int. Conf. Learning Representations (ICLR 2024) (2024).
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., Zhou, M.: MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. In: Advances in Neural Information Processing Systems 33 (NeurIPS 2020), pp. 5776–5788 (2020).
Wilson, E. B.: Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association 22(158), 209–212 (1927). https://doi.org/10.1080/01621459.1927.10502953
Panickssery, A., Bowman, S. R., Feng, S.: LLM evaluators recognize and favor their own generations. In: Advances in Neural Information Processing Systems 37 (NeurIPS 2024) (2024).
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., et al.: Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In: Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track (2023).
Scarlatos, A., Lee, J., Woodhead, S., Lan, A.: Simulated students in tutoring dialogues: Substance or illusion? In: Proc. 64th Annual Meeting of the ACL (ACL 2026), Long Papers, pp. 42349–42385. San Diego, CA, USA (2026). https://doi.org/10.18653/v1/2026.acl-long.1960
Downloads
Posted
Categories
License
Copyright (c) 2026 Research Archive of Rising Scholars

This work is licensed under a Creative Commons Attribution 4.0 International License.