Please use this identifier to cite or link to this item:
https://dspace.iiti.ac.in/handle/123456789/19312
| Title: | Deep learning for mathematical proof synthesis: leveraging leancopilot |
| Authors: | Mishra, Adarsh |
| Supervisors: | Mouli, Sasank |
| Keywords: | Computer Science and Engineering |
| Issue Date: | 20-May-2026 |
| Publisher: | Department of Computer Science and Engineering, IIT Indore |
| Series/Report no.: | MT488; |
| Abstract: | Why is writing proofs in Lean 4 hard? Simple. The system is exact. Lean is expressive enough to formalize a large part of undergraduate and graduate mathematics, and its kernel checks every proof step, so a theorem it compiles has a high correctness guarantee. But to get there, you have to learn Lean syntax, Mathlib naming conventions, proof tactics, and the structure of machine-checkable mathematical reasoning. A proof that is obvious on paper may go through many iterations before Lean accepts it. This thesis studies deep-learning assisted mathematical proof synthesis, especially LeanCopi-lot. LeanCopilot connects a language model to Lean's proof environment, and can suggest tactic sequences for the current proof state [22]. The model is helpful, but it's not automati-cally right on what it suggests. Generated proofs can use non-existent lemmas, apply tactics to goals outside their supported fragment, create type mismatches, or leave goals unresolved. The central argument of this thesis is that these failures are not to be dismissed as noise. They can be gathered, categorized and used to inform improvements to the model. The work has four contributions. First, it gives an end-to-end work flow of Lean 4, Mathlib, and LeanCopilot. Including a sub-problem decomposition strategy to decompose difficult proofs into smaller verified lemmas. Second, it provides a framework to formalize multiple-choice mathematical questions as separate Lean theorems, with the correct option proved and the incorrect options refuted. An example of a JEE style polynomial demonstrates the procedure. Third, it sets up a failure-analysis pipeline that pools the generated Lean files, gathers diagnostic messages and classifies failures into seven categories: type mismatches, hallucinated lemma names, tactic failures, incomplete proofs, syntax or formalization errors, timeouts, and environment errors. Fourth, it employs the resulting taxonomy to inform QLoRA-based fine-tuning on verified theorem-proof pairs. The experimental results show that LLM-assisted proving is most effective when generation is combined with formal verification and systematic feedback. For the reported held-out evaluation, we increased compilation success in percentage from 42 to 71, proof completion from 38 to 68, and decreased hallucinated lemma usage from 31 to 12. The thesis ends with the conclusion that language models should be viewed as proof proposal systems instead of proof authorities: they can reduce the search burden but Lean remains the final correctness checker. Keywords: Lean 4, LeanCopilot, formal verification, automated theorem proving, large language models, QLoRA, failure analysis, MCQ verification, parameter-efficient fine-tuning. |
| URI: | https://dspace.iiti.ac.in/handle/123456789/19312 |
| Type of Material: | Thesis_M.Tech |
| Appears in Collections: | Department of Computer Science and Engineering_ETD |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| MT_488_Adarsh_Mishra_2402101004.pdf | 28.68 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.
Altmetric Badge: