Caltech Mathathon – first hackathon ever devoted to research level mathematics
270 points
• 7 days ago
• Article
Link
California Institute of Technology 将于 2026 年 10 月 30 日至 11 月 1 日举办 Caltech Mathathon,这将成为首个完全以研究级数学为主题的 hackathon,具有重要的里程碑意义。此次活动紧随该领域一系列重大突破之后:AI 已成功攻克诸如距今约 80 年的 Erdos planar unit-distance conjecture 、构造 non-sofic groups,以及在 six-sphere 复杂结构方面取得潜在进展等长期难题。
这些快速进展凸显了关于学科未来的关键问题,尤其是人工智能将如何缩短从初步构想到同行评审发表之间的时间。活动还将探讨:在 AI 日益展现出解决复杂猜想能力的时代,人类数学家的角色将如何演变。组织者希望通过汇集全球数学人才,实时应对这一系列系统性变化。
在 40 小时的挑战赛中,100 支队伍将获得超过 200 万美元的 AI 额度,并可使用多款前沿模型。参赛者将专注于攻克未解的猜想并发展新的数学理论。赛后,各队需向由顶尖数学家组成的评审团展示成果,评审团将评估研究内容的实质价值及参赛者对成果的理解深度。
比赛将设一轮现场奖项,以表彰活动期间展示出的最有前景的研究成果;在更广泛的数学界有足够时间验证这些工作的有效性后,将再颁发第二轮奖项。该活动得到众多科技公司和风险投资公司的支持,作为一次实地试验,Mathathon 旨在检验人类直觉与机器智能在推动纯数学发现方面的协作潜力。
The California Institute of Technology is set to host the Caltech Mathathon from October 30 to November 1, 2026, marking a significant milestone as the first hackathon dedicated entirely to research-level mathematics. This event arrives on the heels of major breakthroughs in the field, where AI has successfully tackled long-standing challenges like the 80-year-old Erdos planar unit-distance conjecture, the construction of non-sofic groups, and potential advancements regarding the complex structure of the six-sphere.
These rapid developments highlight critical questions about the future of the discipline, specifically how artificial intelligence can accelerate the timeline from initial ideation to peer-reviewed publication. The event also seeks to explore the evolving role of the human mathematician in an era where AI is demonstrating an increasing capacity to solve complex conjectures. By bringing together global mathematical talent, the organizers aim to address these systemic shifts in real-time.
During the 40-hour challenge, one hundred teams will be equipped with over 2 million dollars in AI credits and access to frontier models. Participants will focus on solving open conjectures and developing new mathematical theories. Following their work, the teams will present their results to a panel of leading mathematicians, who will evaluate both the substance of their findings and the depth of their comprehension.
The competition will feature an initial round of prizes for the most promising results presented during the event. A subsequent round of awards will be issued once the broader mathematics community has had sufficient time to verify the validity of the work. Supported by a wide array of technology companies and venture firms, the Mathathon serves as a practical experiment to test the collaborative potential of human intuition and machine intelligence in the pursuit of pure mathematical discovery.
105 comments • Comments Link
• 组织者正在举办一场由学生主导的 Mathathon,旨在探索 AI 与研究级数学的交叉领域,强调促进负责任的使用,而非单纯追求速度。
• 批评者质疑 40 小时黑客松形式的必要性,认为当前 AI 在数学上的进展更多依赖长时间运行的自主会话,而不是传统黑客松中那种密集同步的协作。
• 活动聚焦人为因素,特别是选取有影响力问题的能力、发挥数学直觉以及引导 AI agents 完成复杂证明的技巧,而不仅仅依赖提示来快速获得结果。
• 有经验的参与者指出,AI 在构建 Gröbner bases 等机械性任务上表现良好,但在生成新颖想法或新证明方法方面较为吃力,因此人在假设生成和引导逻辑推理中的作用至关重要。
• 有人担心大型 AI labs 正在利用此类活动获取廉价人力,用于验证、清理并以"人类认证"掩饰其不透明且未经证实的模型输出。
• 一些参与者担忧,过度依赖 LLMs 进行数学研究可能削弱基础研究技能,可能导致一代人在没有 AI 辅助时难以深入思考问题。
• 将其称为"首个"研究级数学黑客松的说法受到挑战,批评者指出 Research Collaboration Workshops 和 Sage Math sprints 等已有悠久传统,表明活动组织者与既有数学研究社区存在脱节。
• 对"thinking traces"的透明性仍存疑虑,观点认为这些日志通常经过其他模型的筛选或摘要,给用户一种关于 AI 如何得出结论的虚假理解。
• 在哲学上也有人反对使用前沿模型,指出这些模型由因掠夺性数据抓取、环境影响和计算资源集中而受批评的公司开发,使用这些工具本身即是在纵容其成功。
• 激励机制(提供大量的 token grants)被部分人视为 AI 公司的战略性营销手段,旨在换取公关价值和数据,而非真正追求科学发现。
这次讨论反映了 AI 驱动的数学研究快速转型与学术界传统价值观之间的深刻张力。组织者将此次活动定位为重塑 AI 使用方式、强调人类直觉不可替代的一种尝试;怀疑者则警告称"prompt-engineering"正在取代深度思考,并将该活动视为科技公司精心设计的营销工具。归根结底,分歧在于人类与 AI 的协作究竟代表着赋能发现的新时代,还是由那些更重视产品产出而非科学实质的公司推动的数学严谨性衰退。 • Organizers are hosting a student-led "Mathathon" to explore the intersection of AI and research-level mathematics, aiming to promote responsible usage rather than simply prioritizing speed.
• Critics question the necessity of a 40-hour hackathon format, arguing that current AI mathematical progress relies more on long-running autonomous sessions than on the intensive, synchronous collaboration typical of traditional software hackathons.
• The event focuses on the human element, specifically the ability to select impactful problems, exercise mathematical intuition, and steer AI agents through complex proofs, rather than just prompting for quick results.
• Experienced participants note that while AI excels at rote tasks like constructing Gröbner bases, it struggles with generating novel ideas or new proof methods, making the human's role in hypothesis generation and directing logical flow essential.
• Concerns are raised that large AI labs are leveraging events like this to obtain cheap human labor for validating, cleaning up, and "human-washing" the outputs of their otherwise opaque and unverified models.
• Some participants express apprehension that over-reliance on LLMs for mathematics may erode fundamental research skills, potentially creating a generation of mathematicians unable to think deeply about problems without AI assistance.
• The characterization of this as the "first" hackathon for research-level mathematics is challenged by those who point to long-standing traditions like Research Collaboration Workshops and Sage Math sprints, suggesting a disconnect between the event organizers and the established math research community.
• Skepticism persists regarding the transparency of "thinking traces," with arguments that these logs are often filtered or summarized by other models, providing users with a false sense of understanding regarding how an AI arrived at its conclusions.
• Philosophical objections are raised against the use of frontier models developed by companies criticized for predatory data scraping, environmental impact, and centralizing compute resources, suggesting that engagement with these tools is inherently complicit in their success.
• The incentives structure—offering significant token grants—is viewed by some as a strategic marketing maneuver by AI companies to gain public relations value and data, rather than a genuine pursuit of scientific discovery.
The discussion reflects a deep tension between the rapid, AI-driven transformation of mathematical research and the traditional values of the academic community. While organizers position the event as a way to "reshape" AI use and emphasize the irreplaceable role of human intuition, skeptics warn of "prompt-engineering" replacing deep thought and characterize the event as a sophisticated marketing vehicle for tech companies. Ultimately, the disagreement hinges on whether human-AI collaboration represents an empowering new era of discovery or a degradation of mathematical rigor facilitated by companies prioritizing product output over scientific substance.