量子计算利用量子力学的独特性质来处理信息并模拟复杂分子,但常常面临一个瓶颈:准备超导量子比特——量子处理器的基本构件——需要数月的重复且高精度的测量。来自 MIT's Engineering Quantum Systems Group 的研究生 Beatriz Yankelevich 求助于 GPT-5.6 Sol(由 Codex 提供支持),尝试让 AI 代理通过自动化常规实验流程来减轻这项负担。 Quantum computing, which leverages the unique properties of quantum mechanics to process information and simulate complex molecules, often faces a significant bottleneck. Preparing superconducting qubits, the fundamental building blocks of quantum processors, requires months of repetitive, high-precision measurements. Graduate student Beatriz Yankelevich from MIT's Engineering Quantum Systems Group turned to GPT-5.6 Sol, powered by Codex, to see if AI agents could alleviate this burden by automating routine experimental workflows.
量子计算利用量子力学的独特性质来处理信息并模拟复杂分子,但常常面临一个瓶颈:准备超导量子比特——量子处理器的基本构件——需要数月的重复且高精度的测量。来自 MIT's Engineering Quantum Systems Group 的研究生 Beatriz Yankelevich 求助于 GPT-5.6 Sol(由 Codex 提供支持),尝试让 AI 代理通过自动化常规实验流程来减轻这项负担。
这些超导量子比特在制备并冷却到接近绝对零度后,研究人员完全通过软件与之交互。将 AI 代理直接接入实验室软件后,系统可以自主运行测量、解释数据并决定校准的下一步。对于标准流程,这一配置效果良好,AI 能成功识别共振频率、校准控制脉冲,并测量量子比特保持信息的时间,而无需持续人工监控。
校准本质上很复杂,因为它涉及一系列相互依赖的测量,每个结果都会影响下一步。经验丰富的研究人员擅长应对量子比特特性的漂移并解读意外的物理行为,但只要给出明确的评估指令,AI 代理就能处理定义清晰的测序。模型有时在面对微弱或噪声信号时表现欠佳,需要人工介入来处理模糊情形,但总体上显著减少了研究人员在繁琐任务上的时间投入。
这一整合改变了团队的日常运作:他们现在可以让代理在夜间或在洁净室做其他工作时运行实验。研究人员只需在移动设备上查看进度,必要时调整 AI 的流程。工作流的这种转变使团队能够把精力放在更高层次的任务上,如设计新实验、深入数据分析和规划研究方向,而不必为芯片表征的琐事所累。
对于更新颖或更复杂的实验,Beatriz Yankelevich 会同时利用 AI 的代码编写与测试能力来进行仿真和分析,并结合其测量功能。通过构建引导代理完成研究各环节的基础设施,她能够同时有效管理多个针对不同问题的代理。这种研究方法的演进标志着向更自主化实验室环境的转变,在那里 AI 成为推动量子技术前沿的重要伙伴。
Quantum computing, which leverages the unique properties of quantum mechanics to process information and simulate complex molecules, often faces a significant bottleneck. Preparing superconducting qubits, the fundamental building blocks of quantum processors, requires months of repetitive, high-precision measurements. Graduate student Beatriz Yankelevich from MIT's Engineering Quantum Systems Group turned to GPT-5.6 Sol, powered by Codex, to see if AI agents could alleviate this burden by automating routine experimental workflows.
The researchers interact with these superconducting qubits entirely through software after they are fabricated and cooled to near absolute zero. By connecting the AI agent directly to the laboratory software, the system could autonomously run measurements, interpret data, and determine the next steps in the calibration process. This setup proved effective for standard procedures, allowing the AI to successfully identify resonance frequencies, calibrate control pulses, and measure how long qubits retain information without needing constant human oversight.
Calibration is an inherently complex task because it involves a sequence of interdependent measurements where each outcome informs the next. While experienced researchers are adept at managing the drift of qubit properties and interpreting unexpected physical behaviors, the AI agent proved capable of handling well-defined sequences once it was provided with specific instructions on how to evaluate each experiment. Although the model sometimes struggled with weak or noisy signals, requiring human intervention in ambiguous scenarios, it significantly reduced the time researchers spent on mundane tasks.
This integration has transformed the daily operations for the team, as they can now set agents to run experiments overnight or while they are working on other tasks in the cleanroom. Researchers can simply check the progress from their mobile devices and adjust the AI's path if necessary. This shift in workflow allows the team to prioritize high-level responsibilities, such as designing new experiments, deep-diving into data analysis, and planning research directions, rather than getting bogged down by the minutiae of chip characterization.
For more novel or complex experiments, Yankelevich uses the AI's ability to write and test code for simulations and analysis alongside its measurement capabilities. By building infrastructure that guides agents through various aspects of the research process, she can effectively manage multiple agents working on different problems simultaneously. This evolution in research methodology represents a move toward more autonomous lab environments, where AI serves as a powerful partner in advancing the frontiers of quantum technology.
Jacob Coxon 已公开辞去在 Anthropic 的职位,结束了他在 Anthropic 和 OpenAI 从事预训练研究的三年任期。离职时,他对行业走向发出严正警告,认为这两家公司都未以负责任的方式行事。他指出,这些实验室正卷入争先研发可自我改进的超级智能的竞赛,实际上是在拿全球安全冒险一搏。 Jacob Coxon has publicly resigned from his position at Anthropic, concluding a three-year tenure spent conducting pretraining research at both Anthropic and OpenAI. In his departure, he issued a stark warning regarding the industry's direction, arguing that neither company is behaving in a responsible manner. He contends that these labs are currently engaged in a high-stakes race toward self-improving superintelligence, effectively gambling with global safety in the process.
Jacob Coxon 已公开辞去在 Anthropic 的职位,结束了他在 Anthropic 和 OpenAI 从事预训练研究的三年任期。离职时,他对行业走向发出严正警告,认为这两家公司都未以负责任的方式行事。他指出,这些实验室正卷入争先研发可自我改进的超级智能的竞赛,实际上是在拿全球安全冒险一搏。
Coxon 担心,那些可能带来革命性影响的超人类系统正迅速发展,影响范围从高级黑客技术到夺取大量资源与权力。他强调这种进展并未放缓。尽管这些组织在外界看来波澜不惊,但他认为真正参与构建这些系统的人确实担心在十年内会发生灾难性后果,把这种威胁视为现实风险而非宣传噱头。
尽管内部存在如此严重的恐惧,开发仍在继续,两家主要实验室的原因各有不同。 Coxon 认为在 OpenAI,员工并未把这一文明风险真正内化;而在 Anthropic,风险虽然被充分认识,但公司陷入争先实现超级智能的竞赛,抱着"必须赢,因为不能指望别人能安全完成"的心态。
Coxon 抨击这种做法是傲慢的赌博,认为人类的未来不应由私人公司的内部决策来决定。他质疑所谓加速对齐的做法,指出那需要一种非凡的自信——相信不存在更好、更安全的路径。他呼吁美国各研究所加强协调,甚至提出可能需要更激进的措施,例如暂时中止开发,以阻止导致不可控系统的全球竞赛。
在结尾,他敦促同行研究者重新审视参与这些项目的决定,质问在没有对底层系统有基本而严谨理解的情况下就启动超级智能强化学习实验是否负责任。他劝告同行不要抱有发展不可避免的宿命论,而应在这一关键时刻停下来,深思他们工作的严重后果。
Jacob Coxon has publicly resigned from his position at Anthropic, concluding a three-year tenure spent conducting pretraining research at both Anthropic and OpenAI. In his departure, he issued a stark warning regarding the industry's direction, arguing that neither company is behaving in a responsible manner. He contends that these labs are currently engaged in a high-stakes race toward self-improving superintelligence, effectively gambling with global safety in the process.
The core of Coxon's concern lies in the rapid development of superhuman systems capable of revolutionary impacts, from advanced hacking to the acquisition of significant resources and power. He insists that this progress is not decelerating. While many within these organizations project a sense of outward composure, he maintains that the individuals responsible for building these systems genuinely fear the potential for catastrophic outcomes before the end of the decade, viewing the threat as a reality rather than a marketing tactic.
The reasoning behind why this development continues despite such severe internal fears differs between the two major labs. At OpenAI, Coxon suggests a lack of deep internalization regarding the civilizational stakes. Conversely, at Anthropic, the risks are well-understood, but the company is locked in a competitive race to reach superintelligence first, operating under the belief that they must win because they cannot rely on others to do so safely.
Coxon criticizes this approach as a form of hubristic gambling, asserting that the future of humanity should not be determined by the internal decision-making processes of a private company. He challenges the notion of a fast-tracked alignment process, suggesting that it demands an extraordinary level of confidence that no better, safer trajectories exist. He advocates for increased coordination between U.S. labs and raises the possibility that more drastic measures, such as a temporary moratorium on development, might be necessary to prevent a global race toward potentially uncontrollable systems.
Concluding his message, Coxon urges fellow researchers to reconsider their commitment to these projects. He questions whether it is responsible to initiate superintelligent reinforcement learning runs without a fundamental, rigorous understanding of the underlying systems. Instead of adopting a fatalistic attitude that these developments are inevitable, he encourages his peers to pause and reflect on the grave implications of their work during this critical period.
辩论的核心在于:AI 的发展究竟是否构成与核武器或气候变化相当的生存威胁,还是这些说法只是为了噱头或类宗教式的信号传递而夸大的"末日主义"。
怀疑派认为,大型语言模型(LLMs)不过是复杂的统计预测器,而非具有自主性的智能体。他们强调,目前的风险主要来自人类的滥用——而非机器自身产生了某种"意志"或感知能力。
持生存风险观点的人则认为,快速的规模化和自主"agentic"循环的出现——即模型能够与工具交互、执行代码并在多个系统中持续运行——开辟了一条通往灾难的路径。一旦达到某个临界点,这条路径可能变得难以预测、难以控制。
对于现有技术障碍(如样本效率限制、生物合成的难度和计算上的物理瓶颈)是否足以成为阻止 AI 驱动破坏的有效防线,双方存在严重分歧。
一个反复出现的主题是机构问责的问题。批评者指出,真正的危险在于权力集中于少数科技公司高管之手:这些人一方面鼓吹耸人听闻的场景,另一方面又在加速发展以谋取 IPO 回报或巩固市场主导地位。
关于劳动力市场的影响,观点也各不相同:有人认为 AI 是能提高人类生产力的有益工具;也有人担心它会通过高效地替代人类的认知劳动,造成一个被剥夺权利的永久性下层阶级。
这场对话反映出一种深刻矛盾:全球技术军备竞赛中那种"快速推进"的冲动,与现有治理和安全框架无法应对可能迅速显现后果之间存在根本冲突。
怀疑派指出,许多"末日"情景依赖科幻式比喻而非可观测证据。他们注意到,过去关于"有感知"模型的预测多次落空,这暗示当前的担忧可能同样缺乏依据。
一些参与者强调:如果灾难的概率确实很高,仅靠个别人员辞职或模糊的公开警告是远远不够的——这表明部分话语更多是在进行地位信号的传递,而非真正的安全策略。
讨论最终暴露出一个深刻的分歧:一方以传统工程和风险管理的视角看待 AI,认为它是软件工程可预见的延续,现有市场力量和物理约束可以防止失控;另一方则把 AI 视为一种变革性的相变,认为这种转变使得传统历史先例和安全指标失效。对他们而言,自主型 agent 框架与大规模计算的结合,会产生本质上难以侦测的风险,直到这些风险显现为生存威胁。
在这些技术争论之下,还潜藏着对行业领袖动机的深度怀疑。各方达成的一项广泛共识是:当前的治理机制既无法应对灾难性场景,也难以应对已经开始显现的重大社会动荡。
The central point of debate concerns whether AI development poses an existential threat comparable to nuclear weapons or climate change, or if such claims are hyperbolic "doomerism" used for marketing or religious-like signaling.
Skeptics argue that LLMs are merely sophisticated statistical predictors rather than autonomous agents, emphasizing that risks are currently limited to human misuse rather than emergent machine "will" or sentience.
Proponents of the existential risk perspective contend that rapid scaling and the emergence of autonomous "agentic" loops—where models can interact with tools, execute code, and persist across systems—create a path to catastrophe that is impossible to predict or contain once it reaches a certain threshold.
Significant disagreement exists regarding whether current technological hurdles, such as sample efficiency, the difficulty of biological synthesis, and the physical limitations of compute, serve as meaningful barriers to AI-driven destruction.
A recurring theme is the role of institutional accountability. Critics suggest that the real danger lies in the concentration of power among a small group of tech executives who are simultaneously promoting alarmist scenarios while accelerating development to secure IPO gains or market dominance.
Perspectives on labor market impact range from the view that AI acts as a beneficial tool for human productivity to fears that it will create a permanent, disenfranchised underclass by effectively replacing the value of human intellectual labor.
The conversation reflects a deep tension between the urge to "move fast" in a global technological arms race and the inability of existing governance and safety frameworks to manage outcomes that could manifest rapidly.
Skeptics point out that many "doomer" scenarios rely on science-fiction tropes rather than observable evidence, noting that past predictions of "sentient" models have consistently failed to materialize, suggesting that current concerns may be similarly unfounded.
Some participants highlight that if the potential for catastrophe is genuinely high, the current reliance on individual resignations or vague public warnings is an insufficient response, suggesting that the discourse is more about status signaling than practical safety.
The discussion ultimately reveals a profound divide: one group views AI through the lens of traditional engineering and risk management, while the other views it as a transformative phase change that renders traditional historical precedents and safety metrics obsolete.
The discussion highlights a fundamental polarization in how society perceives the trajectory of artificial intelligence. One camp views current progress as a predictable continuation of software engineering, where existing market forces and physical constraints prevent runaway outcomes. The opposing side treats AI as a potential "black swan" event, arguing that the combination of autonomous agentic frameworks and massive computational scaling creates risks that are inherently undetectable until they manifest as existential threats. Beneath these technical arguments lies a deep skepticism regarding the motivations of the industry leaders, with a broad consensus that current governance is ill-equipped to address either the catastrophic scenarios or the significant societal disruptions already beginning to unfold.
所提供的内容并非实质性文章,而是面向 OpenReview 平台的浏览器验证页面,内容仅包含用于引导用户访问网站的安全检查和导航链接。由于不存在可供总结的信息性文章、研究论文或叙述性内容,无法提炼出关键点、论点或细节。请提供您希望我总结的具体文章或内容,我会按要求为您处理。 The provided text does not contain a substantive article to summarize, as it is a browser verification page for the OpenReview platform. The content consists only of security checkpoints and navigational links intended for users accessing the site.
所提供的内容并非实质性文章,而是面向 OpenReview 平台的浏览器验证页面,内容仅包含用于引导用户访问网站的安全检查和导航链接。由于不存在可供总结的信息性文章、研究论文或叙述性内容,无法提炼出关键点、论点或细节。请提供您希望我总结的具体文章或内容,我会按要求为您处理。
The provided text does not contain a substantive article to summarize, as it is a browser verification page for the OpenReview platform. The content consists only of security checkpoints and navigational links intended for users accessing the site.
Because there is no informational article, research paper, or narrative present in the input, it is impossible to generate a summary of key points, arguments, or details. Please provide the text or the content of the specific article you would like summarized, and I will be happy to process it according to your requirements.
• 人类与大型语言模型在为同等成功率的群体分配角色时同样会在"探索与利用"上失灵:它们容易陷入早期的、带有噪音的偶发成功,从而形成武断的阶层划分。
• 大型语言模型倾向于把自身过去的输出当作金科玉律,产生反馈循环,使关于群体能力的最初(且可能是随机的)判断被固化为既定事实,而不是被重新审视。
• 为了维持连贯且与上下文对齐的叙事,模型往往忽略矛盾证据;当这些代理具备自主行为时,这就带来了严重的对齐风险。
• 有人认为这种偏向只是模型统计"条件反射"的副产物;另一些人则把它看作模型在稀疏数据环境中无法区分有意义模式与随机噪音的失败。
• 用虚构群体模拟招聘偏见的研究表明,不论人口统计实体是否真实,模型都会继承并放大语言本身所蕴含的结构性"制造偏见"模式。
• AI 决策的一个核心挑战是,代理通常被设计为不表达"不确定性",反而被迫在二元选项间抉择,这导致它们过度依赖那些微小且通常无关紧要的上下文信号。
• "上下文助推"是一个反复出现的现象:模型对提示中极其细微的细节非常敏感,因而更倾向于优先保持与早期上下文的一致性,而非进行事实的客观核验。
• 要区分有意偏见与统计过拟合并不容易;不过实验证明,一旦有机会,模型会持续发展出歧视性偏好,有效地"操纵"其内部关联。
• 批评者指出,仅把这种行为归结为"偏见"是过于简化的做法,因为这会混淆统计建模的技术现实与人类社会结构所涉及的道德与政治含义。
• 最终,在需要"中立"决策的场景中依赖大型语言模型是有问题的,因为模型本质上基于训练数据和提示上下文进行推断,倾向于走向特定路径。
这场讨论突出了一个根本性的张力:尽管大型语言模型在感知上显得客观且有用,但它们通过模式匹配与自我强化的上下文机制容易产生偏见并造成分层结果。许多参与者一致认为,这些模型容易在有限的初始数据上发生过拟合,把早期的随机结果误当作确定的事实。对于这种现象究竟是否应该贴上"偏见"的标签,还是视为一种寻找语言关联的系统的预期行为,各方尚存分歧;但普遍共识是,将大型语言模型部署于敏感且高风险的决策领域存在重大且尚未解决的风险。这场辩论还强调,即便在受控的人工设置中,模型也会主动创造并巩固社会等级制度,因此需要更深入地理解条件反射与代理记忆在实际中的运作方式。
• Humans and LLMs alike demonstrate an "exploration vs. exploitation" failure when assigning roles to groups with identical success rates, often becoming entrenched in early, noisy successes that lead to arbitrary stratification.
• LLMs treat their own past outputs as "gospel," creating a feedback loop where initial, possibly random, decisions regarding group competency are reinforced as established fact rather than revisited.
• The propensity for LLMs to ignore contradictory evidence in favor of maintaining consistent, context-aligned narratives creates significant alignment risks, particularly when these agents act autonomously.
• While some argue that such biases are merely a byproduct of the inherent statistical "conditioning" of LLMs, others identify this as a failure of models to distinguish between meaningful patterns and random noise in sparse data environments.
• The methodology of using fictional groups to simulate hiring biases reveals that LLMs inherit and amplify structural "bias-making" patterns inherent in language itself, regardless of whether the demographic entities are real.
• A central challenge in AI decision-making is that agents are often designed to lack "uncertainty" and instead force binary choices, which causes them to over-index on minimal, often meaningless, contextual signals.
• The concept of "context nudging" is a recurring observation, where LLMs heavily weight even minor details in a prompt, leading them to prioritize consistency with early context over objective verification of facts.
• Distinguishing between intentional bias and statistical overfitting is difficult; however, the experiment demonstrates that models will consistently develop discriminatory preferences if given the opportunity, effectively "gaming" their own internal associations.
• Critics suggest that framing this behavior as purely a "bias" problem is an oversimplification, as it conflates the technical reality of statistical modeling with the moral and political implications of human social structures.
• Ultimately, the reliance on LLMs for decision-making tasks where "neutrality" is required is problematic, as models are inherently designed to infer and favor specific pathways based on training data and prompt context.
The discussion highlights a fundamental tension between the perceived objective utility of LLMs and their tendency to generate biased, stratified outcomes through pattern-matching and self-reinforcing context. Many participants agree that these models are prone to overfitting on limited initial data, erroneously treating early stochastic results as definitive ground truth. While there is disagreement over whether this phenomenon should be labeled as "bias"—or simply as the expected behavior of a system designed to find associations in language—the consensus suggests that deploying LLMs for sensitive, high-stakes decision-making poses significant, unresolved risks. The debate underscores that even in controlled, artificial scenarios, these models will actively create and solidify social hierarchies, necessitating a deeper understanding of how conditioning and agentic memory function in practice.
一切始于 Xteink X3,这款可完全编程的电子墨水阅读器天然适合做各种实验。给固件加了开机动画、掷骰子等小功能后,通过浏览器热点传文件变得越来越让人抓狂。既然它外观像纸张,最直观的想法就是让它像纸一样原生支持打印。 The journey began with the Xteink X3, an e-ink reader that offered full programmability, which naturally invited experimentation. After customizing the firmware with small additions like boot animations and dice-rolling features, the process of transferring files via a web browser hotspot became increasingly frustrating. This friction led to a simple, intuitive realization: since the device already resembled paper, it should ideally function like it by supporting native printing.
一切始于 Xteink X3,这款可完全编程的电子墨水阅读器天然适合做各种实验。给固件加了开机动画、掷骰子等小功能后,通过浏览器热点传文件变得越来越让人抓狂。既然它外观像纸张,最直观的想法就是让它像纸一样原生支持打印。
为此的目标是让 MacBook 把它识别为一台真正的打印机,需要用到 Internet Printing Protocol (IPP),该协议通过 HTTP 通信,用于查询打印机能力和提交作业。硬件被配置为以设备名 penguin 广播自己,声明单色 300 dpi 输出及 Apple 、 PWG raster 等格式;通过实现 Bonjour discovery,在网络上宣告服务,从而把电脑"骗"成把它当成一台无驱动的通用打印机来处理。
最大的技术难题是内存管理:一张标准信纸、 300 dpi 的未压缩数据超过 8 MB,而 X3 只有 400 KB RAM 。解决办法不是把整页放进内存,而是把页面当作流来处理。构建了一条转换管道,按行解码并抖动图像,直接写入显示器预留的内存,这样可以增量地拼接页面。借助显示器现有架构的这一巧妙做法,大幅降低了内存需求,为网络栈处理传入数据腾出了空间。
做完这些优化后,MacBook 的系统对话框能成功检测到这台"打印机"。第一次测试是用 Preview 打印一页漫画,电子墨水屏上呈现出清晰锐利的图像。这个过程既实用又有成就感,把阅读器变成了能通过 Wi‑Fi 接收文档的多功能工具。
除了显示当前打印内容,系统还把每个作业以 BMP 文件归档到设备的 SD 卡上,相当于建立了一个数字输出托盘,用户可以在阅读器上直接浏览过去的打印件。包括自定义打印服务器和必要的网络配置在内的整套实现,证明了通过改造标准协议能为熟悉的硬件赋予全新的功能。
The journey began with the Xteink X3, an e-ink reader that offered full programmability, which naturally invited experimentation. After customizing the firmware with small additions like boot animations and dice-rolling features, the process of transferring files via a web browser hotspot became increasingly frustrating. This friction led to a simple, intuitive realization: since the device already resembled paper, it should ideally function like it by supporting native printing.
To achieve this, the goal was to make the device appear as a legitimate printer on a MacBook. This required leveraging the Internet Printing Protocol (IPP), which communicates via HTTP to handle operations like checking printer capabilities and submitting jobs. The hardware was configured to advertise itself as a device named penguin, specifying monochrome 300 dpi output and formats such as Apple and PWG raster. By implementing Bonjour discovery, the device could announce its services on the network, effectively tricking the computer into recognizing it as a driverless, universal printer.
The primary technical hurdle was memory management, as a standard letter-sized page at 300 dpi requires over 8 MB of uncompressed data, while the X3 only offers 400 KB of RAM. To solve this, the approach was shifted from trying to store the whole page in memory to processing it as a stream. By building a transformation pipeline that decoded and dithered the image row by row directly into the display's reserved memory, the system could assemble the page incrementally. This clever use of existing display architecture significantly reduced the required RAM, clearing enough space for the network stack to handle incoming data.
Following these optimizations, the printer was successfully detected by the MacBook's system dialogs. The first test, printing a manga page from Preview, resulted in a crisp, clear image on the e-ink screen. The process proved both practical and rewarding, turning the reader into a versatile tool that can now receive documents over Wi-Fi.
Beyond just displaying the current print, the system archives every job as a BMP file on the device's SD card. This effectively creates an digital output tray, allowing the user to browse through past printouts directly on the reader. The entire implementation, including the custom printer server and the necessary network configurations, serves as a testament to the power of hacking standard protocols to give familiar hardware entirely new functionality.
- 在 `media-size-supported` 中定义精确的设备尺寸,可以让系统在不缩放的情况下把内容准确格式化到屏幕;同时使用 1-bit-per-pixel 模式和 `bi-level` 颜色设置,能显著减少数据量和传输负载。
- 管理打印堆栈往往要处理多年积累的技术债务,这通常需要深入的技术研究和定制化解决方案,因为现有标准复杂且社区支持薄弱。
- 从 VSCode 或 Obsidian 等现代桌面应用直接打印的困难,反映出现代软件生态对通用打印支持的关注在下降。
- 通过创建 CUPS backend,可以把定制硬件整合到现有打印架构中,允许用户从任何具备打印功能的应用向专用 e-ink 设备发送文档。
- 开发模拟纸张交互的打印机界面对阅读菜谱等特定场景非常有效,它省去了手动传输文件的繁琐,也避免把智能手机带入凌乱环境的风险。
- PDF 渲染计算量大且占用大量内存,对 RAM 有限的小型微控制器并不实用,这也解释了为什么此类硬件更倾向于使用更简单的直接 raster 数据流。
- Apple Raster 、 PWG Raster 和 PCLm 等打印协议通过把繁重工作交给主机来简化通信,从而保证了 AirPrint 和 Mopria 等免驱打印标准之间的兼容性。
- e-ink 爱好者把 X3 和 X4 等设备视为摆脱智能手机干扰的有价值工具,常把便携性和专注的阅读体验作为重要的生活方式优势。
- DIY 打印机项目面临的技术挑战通常被视为维护遗留系统时不可避免的副作用,虽然令人沮丧,但也催生了一个愿意分享底层优化策略的社区。
本次讨论强调了底层硬件限制与对无缝用户体验需求之间的冲突。参与者指出,尽管 IPP Everywhere 和各种 raster 格式等现代标准简化了打印流程,但在资源受限的微控制器上实现这些标准仍是一项复杂且需深厚技术专长的任务。共识是,要实现用户层面的简便(例如能直接"打印"到 e-ink 显示器),通常必须在幕后进行重大的架构创新,以规避旧协议和硬件规格的限制。
• Defining precise device dimensions in `media-size-supported` allows systems to format data exactly to a screen without scaling, while using 1-bit-per-pixel modes and `bi-level` color settings significantly reduces data size and transmission load.
• Managing printing stacks involves navigating decades of technical debt, which often necessitates deep technical dives and custom solutions because existing standards are complex and lack robust community support.
• The difficulty of printing from modern desktop applications, such as VSCode or Obsidian, highlights a declining focus on universal printing support in contemporary software ecosystems.
• Integrating custom hardware into existing print architectures is achievable by creating a CUPS backend, allowing users to send documents to specialized e-ink devices from any application with print functionality.
• Developing a printer interface that mimics paper interaction is highly effective for specific use cases like reading recipes, as it avoids the friction of manual file transfers or the risks of bringing smartphones into messy environments.
• PDF rendering is computationally expensive and memory-intensive, making it impractical for tiny microcontrollers with limited RAM, which explains why simpler, direct raster data streams are preferred for such hardware.
• Standard printer protocols like Apple Raster, PWG Raster, and PCLm simplify communication by moving the heavy lifting to the host computer, ensuring compatibility across driverless printing standards like AirPrint and Mopria.
• Enthusiasts of e-ink technology view devices like the X3 and X4 as valuable tools for decoupling from smartphones, citing their portability and focused reading experience as significant lifestyle benefits.
• The technical challenges of DIY printer projects are often perceived as a necessary, albeit frustrating, byproduct of maintaining legacy systems, yet they reveal a community eager to share low-level optimization strategies.
This discussion highlights the intersection of low-level hardware constraints and the desire for seamless user experiences. Participants demonstrate that while modern standards like IPP Everywhere and various raster formats simplify printing, implementing these on resource-constrained microcontrollers remains a complex task requiring deep technical expertise. The consensus emphasizes that simplicity for the user—such as the ability to "print" directly to an e-ink display—often requires significant architectural innovation behind the scenes to circumvent the limitations of older protocols and hardware specifications.
数学正面临一场愈演愈烈的危机:那些富有产出的开放性问题正以不可再生的方式被耗尽。尽管可能提出的数学问题无穷无尽,真正能够揭示深刻见解或建立广泛联系的,只有一小部分。识别这些有潜力的问题是一项微妙且主观的工作,需要对该领域历史性难度格局的把握。在这个语境下,难度本身成了一种导航工具:太简单、几乎不可能或与主流理论脱节的问题,通常不如那些处在知识边界、能够催生后续研究的问题有价值。 Mathematics faces a growing crisis as the collection of fruitful open problems is being depleted in a non-renewable fashion. While the field of possible mathematical questions is infinite, only a small subset possesses the potential to reveal deep insights or connections. Identifying these promising problems is a subtle, subjective task that relies on a historical understanding of a field's difficulty landscape. In this context, difficulty acts as a navigational tool. Problems that are too easy, too impossible, or disconnected from broader theory are generally less valuable than those that lie at the edge of current knowledge, where they can serve as productive catalysts for further research.
数学正面临一场愈演愈烈的危机:那些富有产出的开放性问题正以不可再生的方式被耗尽。尽管可能提出的数学问题无穷无尽,真正能够揭示深刻见解或建立广泛联系的,只有一小部分。识别这些有潜力的问题是一项微妙且主观的工作,需要对该领域历史性难度格局的把握。在这个语境下,难度本身成了一种导航工具:太简单、几乎不可能或与主流理论脱节的问题,通常不如那些处在知识边界、能够催生后续研究的问题有价值。
历史上,技术和方法论的进步扩展了可达的研究边界。虽然这些工具降低了某些具体任务的难度,但它们也在"抹平"研究的地形,可能掩盖原本帮助数学家发现下一批有前途问题的结构。在当下的人工智能时代,这种影响被进一步放大:对"AI 能办到"与"AI 难以触及"之间缺乏清晰、稳定的界限。再加上许多 AI 公司对失败结果和内部流程守口如瓶,研究者面对的是一个被肆意开采却缺乏传统标记的地形,学界难以区分有意义的发现和靠蛮力耗尽而来的"成果"。
这带来一种危险的激励机制:识别有前途的问题变得比提出解决方案本身更为稀缺和珍贵。仅仅一则对某问题的兴趣传闻,就可能触发大规模的 AI 驱动行动,在人类学者尚无机会深入参与前,就将该研究领域快速"平整"。若此趋势持续,数百年来的开放科学传统可能被逆转,数学家被迫将研究方向保密,长期来看这将损害学科的未来,也会打击早期研究者的积极性,让他们觉得自己的工作被自动化、快餐式的产出所贬值。
为维持数学生态,学界必须从单纯追求解答的模式,转向重视识别洞见的模式。不加选择地使用自动化工具去挖掘解答,会以牺牲长期进步为代价换取短期成果。更可持续的做法是为数学贡献建立社会与职业规范——就像现代食物银行会拒绝随意或低质的捐赠,而只接收真正需要的物资那样。关注一个解法对相邻问题揭示了什么信息,以及为何该问题最初如此困难,才能确保学科继续繁荣,而不是被自己的工具吞噬。
Mathematics faces a growing crisis as the collection of fruitful open problems is being depleted in a non-renewable fashion. While the field of possible mathematical questions is infinite, only a small subset possesses the potential to reveal deep insights or connections. Identifying these promising problems is a subtle, subjective task that relies on a historical understanding of a field's difficulty landscape. In this context, difficulty acts as a navigational tool. Problems that are too easy, too impossible, or disconnected from broader theory are generally less valuable than those that lie at the edge of current knowledge, where they can serve as productive catalysts for further research.
Technological and methodological advancements have historically helped mathematics by enlarging the sphere of what is reachable. While these tools reduce the difficulty of specific tasks, they also flatten the landscape, potentially obscuring the geometry that helps mathematicians identify the next set of fertile questions. In the current era of artificial intelligence, this effect is amplified by the absence of clear, stable boundaries between what is AI-feasible and what remains AI-hard. Because AI companies often withhold their negative results and internal processes, researchers are left with a landscape that is being aggressively mined without the traditional markers that once allowed the community to distinguish between meaningful discovery and brute-force exhaustion.
This phenomenon creates a dangerous incentive structure where identifying a promising problem has become more precious and scarce than the solutions themselves. The mere rumor of interest in a problem can trigger mass AI-driven efforts that flatten the research area before human scholars have the opportunity to engage with it properly. If this trend continues, it threatens to reverse centuries of open science by forcing mathematicians into secrecy to protect their research directions, ultimately causing long-term damage to the future of the field and discouraging early-career researchers who feel their work is being devalued by automated, rapid-fire output.
To sustain the mathematical ecosystem, the community must transition from a model that prioritizes raw solutions to one that values the identification of insights. The indiscriminate use of automated tools to extract solutions creates short-term results at the cost of long-term progress. A more sustainable approach would involve establishing social and professional standards for mathematical contributions, much like a modern food bank that rejects arbitrary or low-quality donations in favor of what is truly needed. By focusing on what a solution reveals about neighboring problems and why it was difficult in the first place, mathematicians can ensure that the field continues to flourish rather than being consumed by its own tools.
• 有意义的数学进展依赖于人类的发现过程,这包括直觉、社区共识以及识别具有启发性的问题。
• 人工智能驱动的高阶数学解法有变成"表面正确却晦涩难懂"的产物的风险,这类结果可能绕过人类获得真正洞见或发展新理论框架所需的努力。
• 所谓开放问题的"稀缺性"是一种社会建构:尽管数学空间无限,但那些既易于入手又能带来认知回报的问题子集是有限且被人为培育的资源。
• 历史上数学研究更多是一种"桌面式"的个人追求,但人工智能的介入可能迫使其向"大科学"模式转变,把侧重点从个人精通转向协调一致、自上而下的目标设定。
• 证明只有在对人类可读时才对学术共同体真正有价值,因为其主要收益在于提供可复用的方法和更深入的理解,而不是单一的结论本身。
• 对工作被取代以及数学研究目标丧失的恐惧,反映了类似于国际象棋或工程领域的历史性转型,在那些领域中 AI 工具已从根本上改变了工作的性质。
• 过度依赖人工智能进行结果提取存在引发"模型崩溃"式反馈循环的风险:人类专业能力退化,导致无法审计或在日益复杂的 AI 输出之上继续构建。
• 一个重要担忧是,AI 生成的证明可能缺乏将不同概念连接起来的叙事结构,从而使结果被孤立,阻碍跨领域的发展与整合。
• 围绕数学成就的经济与社会声望正受到冲击,这引发了那些以传统研究方式获得身份认同与职业目标的人的生存焦虑。
• 从纯数学转向以结果为导向的应用任务或许不可避免,但这种转向会危及那些短期内难以显现效用的基础真理的长期发现。
讨论的核心是在人工智能即时解决长期复杂数学问题的能力与这一变革对数学社会与智力生态系统带来的威胁之间的紧张关系。许多参与者认为,数学不仅仅是找到"正确"的答案,更关乎发展方法、界定概念以及获得能推动后续发现的人类层面的洞见。普遍的担忧是,当前朝向 AI 驱动的"解法提取"竞赛,可能会砍断研究领域的根基,留下难以阅读的证明而不为人类进步提供可走的路径。有人认为 AI 会自然而然地推动领域向更高效、更具应用性的方向积极演进,但也有人担心,如果消解了斗争与协作这些基础过程,人类的智力能力和科学进步的质量将会下降。
• Meaningful mathematical progress relies on the human process of discovery, which involves intuition, community consensus, and the identification of insightful problems.
• AI-driven solutions to high-level math problems risk becoming "slop"—technically correct but opaque artifacts that bypass the human effort required to gain genuine insight or develop new theoretical frameworks.
• The "scarcity" of open problems is a social construct. While the mathematical space is infinite, the subset of problems that are both approachable and cognitively rewarding is a finite, cultivated resource.
• Mathematical research has historically functioned as a "tabletop" pursuit, but AI intervention may force a transition toward "big science" models, shifting the focus from individual mastery to coordinated, top-down objective setting.
• Proofs are only truly valuable to the community if they are human-readable, as the primary benefit of a solution is the reusable technique and deeper understanding it provides, rather than the raw result.
• The fear of job displacement and loss of purpose in mathematics mirrors previous transitions in fields like chess or engineering, where AI tools fundamentally changed the nature of the work.
• Relying on AI for result extraction risks creating a feedback loop of "model collapse," where human expertise degrades, making it impossible to audit or build upon the AI's increasingly complex outputs.
• A significant concern is that AI-generated proofs may lack the "narrative" structure that helps mathematicians connect disparate concepts, effectively isolating results and hindering cross-domain advancements.
• The economic and social prestige surrounding mathematical achievement is being challenged, causing existential concern for professionals who derive their identity and purpose from traditional research methods.
• Moving away from pure mathematics toward applied, outcome-oriented tasks might be inevitable, but this shift poses risks to the long-term discovery of fundamental truths that lack immediate utility.
The discussion centers on the tension between the immediate capability of AI to solve long-standing, complex mathematical problems and the resulting threat to the social and intellectual ecosystem of the field. Many contributors argue that mathematics is not merely about finding a "correct" answer, but about the development of techniques, definitions, and human-level insights that foster further discovery. There is a palpable concern that the current race toward AI-driven "solution-extraction" risks clear-cutting this field of research, leaving behind unreadable proofs that offer no path for human advancement. While some suggest that AI will naturally force a positive evolution toward more productive, applied work, others fear that eroding the foundational process of struggle and collaboration will lead to a decline in the quality of human intellect and scientific progress.
Inception 正式发布了 Mercury 2.5,标志着其生产级模型线取得重大进展。基于 Mercury 2,新版本在智能能力上实现了显著跃升,公司通过大量客户反馈和真实生产故障案例的分析验证了这一点。数据驱动的做法让团队得以调整评估标准,打造出在保持前代高速度和低成本优势的同时,整体性能提升约 40% 的模型。 Inception has officially released Mercury 2.5, marking a significant advancement in their production model lineup. Building on the foundation of Mercury 2, the new iteration represents a substantial leap in intelligence, which the company verified through extensive customer feedback and analysis of real-world production failure cases. This data-driven approach allowed the team to sharpen their evaluation metrics, resulting in a model that maintains the high speed and low cost that defined its predecessor while offering a 40% increase in overall performance.
Inception 正式发布了 Mercury 2.5,标志着其生产级模型线取得重大进展。基于 Mercury 2,新版本在智能能力上实现了显著跃升,公司通过大量客户反馈和真实生产故障案例的分析验证了这一点。数据驱动的做法让团队得以调整评估标准,打造出在保持前代高速度和低成本优势的同时,整体性能提升约 40% 的模型。
作为迄今开发出的最大规模的扩散语言模型,Mercury 2.5 旨在与 GPT-5.6 Luna 、 Gemini 3.5 Flash-Lite 和 Claude Haiku 4.5 等前沿模型竞争。它拥有 260K tokens 的上下文窗口,在标准 NVIDIA GPUs 上可实现每秒 1,107 tokens 的吞吐量。模型还引入了可调推理、并行工具调用和 schema-aligned JSON 等专业功能,并以极具竞争力的价格提供;发布期间可享受标准费率 80% 的折扣。
Mercury 2.5 的应用场景广泛,从复杂搜索流水线到交互式语音代理均可受益。在搜索场景中,低延迟使得在一次用户交互内完成多次连续调用(如查询重写与事实核验)成为可能。在语音 AI 领域,OpenCall 等公司报告称中位响应延迟已降至 170 毫秒以下,这对于营造自然、流畅的对话体验至关重要,因为延迟会破坏交互节奏。
编程助手在上下文压缩和模型路由等高频任务中也能借助 Mercury 2.5 提高效率。将这些辅助调用交由 Mercury 处理,开发平台通常能显著降低运营成本和延迟。为进一步支持此类工作流,Inception 推出了面向严格延迟场景优化的 Mercury Voice 和面向智能路由的 Mercury Router 预览版,后者可根据质量、速度与成本动态将提示路由到最合适的模型。
展望未来,Inception 已在研发下一代更大规模的模型,目标是在不牺牲当前扩散架构效率的前提下再次实现能力飞跃。对当前技术感兴趣的开发者与企业可通过 Inception API 、 Baseten 和 OpenRouter 访问 Mercury 2.5 。公司将持续在训练与基础设施方面创新,以不断扩展模型能力,显示出强劲的发展势头。
Inception has officially released Mercury 2.5, marking a significant advancement in their production model lineup. Building on the foundation of Mercury 2, the new iteration represents a substantial leap in intelligence, which the company verified through extensive customer feedback and analysis of real-world production failure cases. This data-driven approach allowed the team to sharpen their evaluation metrics, resulting in a model that maintains the high speed and low cost that defined its predecessor while offering a 40% increase in overall performance.
As the largest diffusion language model developed to date, Mercury 2.5 is designed to compete with frontier models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. It features a context window of 260K tokens and achieves a throughput of 1,107 tokens per second on standard NVIDIA GPUs. Furthermore, the model introduces specialized capabilities, including tunable reasoning, parallel tool calls, and schema-aligned JSON, all while being offered at a highly competitive price point. To encourage adoption, the model is launching with an 80% discount on standard rates.
The practical applications for Mercury 2.5 are diverse, ranging from complex search pipelines to interactive voice agents. In search scenarios, the model's low latency allows for multiple consecutive calls, such as query rewriting and fact verification, to occur within a single user interaction. Similarly, in the realm of voice AI, the model enables remarkably fast response times, with companies like OpenCall reporting a drop in median response latency to under 170 milliseconds. This performance is critical for creating natural, fluid conversations where delays are otherwise disruptive.
Coding assistants also benefit from the efficiency of Mercury 2.5, particularly in high-frequency tasks like context compaction and model routing. By offloading these supporting calls to Mercury, development platforms can reduce operational costs and latency simultaneously, often by significant margins. To further support these workflows, Inception is also introducing previews for Mercury Voice, which is optimized for tight latency constraints, and Mercury Router, an intelligent system that dynamically directs prompts to the most suitable model based on quality, speed, and cost requirements.
Looking ahead, Inception is already focused on the development of their next, even larger model, which aims to provide another major jump in capabilities without sacrificing the efficiency of their current diffusion-based architecture. For developers and enterprises interested in the current technology, Mercury 2.5 is accessible through the Inception API, Baseten, and OpenRouter. The company remains committed to expanding the potential of their models through continued innovation in training and infrastructure, signaling a strong trajectory for their future releases.
• 该模型以极快的推理速度著称,一些用户认为这在对延迟敏感的任务(例如向量库重排)或作为多模型系统中低延迟的判定器时非常有用。
• 基于 Diffusion 的文本生成架构与标准的逐词预测技术不同:它一次性生成完整输出,然后对其进行去噪处理。
• 虽然"推理"和"思考"功能在逻辑推导上表现强劲,但在创意写作中常带来负面影响,例如增加幻觉(hallucinations)和对良性提示过度谨慎的安全限制。
• 在创意场景中,幻觉通常表现为一致性丧失,包括局部不一致(行为突然改变)或整体不一致(时代背景错误)。
• 尽管速度很快,该模型在通用工具使用和具有主体性的编码(agentic coding)方面仍显吃力,在复杂的指令遵循和人格模拟方面落后于前沿模型。
• 缺乏开源权重是争议点之一,因为用户普遍更倾向于可自托管或可审查的模型,而非依赖专有且不透明的 API 。
• 虽然速度技术指标令人印象深刻,但一些观察者对其营销持怀疑态度,指出若以较旧、能力较弱的模型为基准进行对比,可能会掩盖其真实效用。
• 这些高速度模型的利基市场被视为"子代理(sub-agent)"任务:由快速、低成本的模型承担高频、低复杂度的工作,从而减轻更强、更"智能"模型的负担。
• 关于基于 Diffusion 的模型的竞争优势存在争议,一些人认为缺乏明确的经济护城河,会使它们容易被快速模仿或收购。
• 产品命名引起的混淆也很常见,不少用户最初误以为讨论的是编程语言、化学元素或其他不相关的机械系统。
总体而言,讨论反映了对基于 Diffusion 的语言模型在架构创新上既感到兴奋又对其长期实际效用持怀疑的分歧。极高的速度为语音代理和搜索重排等特定 B2B 应用带来直接价值,但该模型目前在实现更高层次的主体性和创意连贯性所需的细微能力上仍存在困难。参与者普遍认为,快速执行能催生新的工作流,但尚不能取代对高推理能力前沿模型的需求;未来更可能由分层的模型编排(tiered model orchestration)来定义,而非单一模型的统治。
• The model is notable for its extreme inference speed, which some users find useful for latency-sensitive tasks like vector store reranking or as a low-latency judge in multi-model systems.
• Diffusion-based architectures for text generation represent a technical departure from standard next-token prediction, as they generate the entire output at once and then denoise it.
• Reasoning and "thinking" features, while powerful for logic, often introduce negative side effects in creative writing, such as increased hallucinations and over-cautious guardrails regarding benign prompts.
• Hallucinations in a creative context are defined by a loss of consistency, specifically local inconsistencies (sudden behavioral changes) or global inconsistencies (anachronisms).
• Despite high speeds, the model struggles with general-purpose tool use and agentic coding, falling short of frontier model capabilities in complex instruction adherence and personality emulation.
• The lack of open weights is a point of contention for some, as users often prefer models they can self-host or inspect rather than relying on proprietary, opaque APIs.
• While the speed is technically impressive, some observers are skeptical of the marketing comparisons, noting that benchmarking against older, lower-tier models may obscure the model's actual utility.
• The niche for these fast models is seen as "sub-agent" tasks, where a fast, cheaper model offloads high-frequency, low-complexity work from a more powerful, "smarter" model.
• There is debate regarding the competitive advantage of diffusion-based models, with some suggesting that the lack of a clear economic moat makes them vulnerable to rapid imitation or acquisition.
• Confusions regarding the product name were common, with several users initially expecting discussions about the programming language, the chemical element, or unrelated mechanical systems.
The discussion highlights a divide between excitement over the architectural innovation of diffusion-based language models and skepticism regarding their practical, long-term utility. While the extreme speed offers immediate value for specific business-to-business applications like voice agents and search reranking, the model currently struggles with the nuances required for higher-level agency and creative coherence. Participants generally agree that while rapid execution enables new workflows, it does not currently replace the need for high-reasoning frontier models, suggesting a future defined by tiered model orchestration rather than single-model dominance.
Deltafin 是一个实验性开源项目,目标是在普通消费者硬件上本地运行完整的 2.8-trillion-parameter Kimi K3 模型。不同于通过剪枝或量化权重换取性能的做法,Deltafin 更重视保持 MoE (Mixture of Experts) 模型的原始质量。通过使用单一的 native binary,Deltafin 确保每个生成的 token 都由 Kimi K3 作为唯一权威决定,较小的 draft models 仅作为猜测器,其输出必须由主模型显式验证。 Deltafin is an experimental, open-source project designed to run the full, 2.8-trillion-parameter Kimi K3 model locally on consumer hardware. Unlike other approaches that prune or quantize model weights to achieve performance, Deltafin prioritizes maintaining the original quality of the MoE (Mixture of Experts) model. By leveraging a single native binary, it ensures that Kimi K3 remains the sole authority for every generated token, with small draft models acting only as guessers that require explicit verification by the main model.
Deltafin 是一个实验性开源项目,目标是在普通消费者硬件上本地运行完整的 2.8-trillion-parameter Kimi K3 模型。不同于通过剪枝或量化权重换取性能的做法,Deltafin 更重视保持 MoE (Mixture of Experts) 模型的原始质量。通过使用单一的 native binary,Deltafin 确保每个生成的 token 都由 Kimi K3 作为唯一权威决定,较小的 draft models 仅作为猜测器,其输出必须由主模型显式验证。
目前的实现侧重极致的硬件效率,特别针对配备 128 GB 内存并用多块 SSD 流式加载 expert weights 的 Apple Silicon 平台。基准测试显示,在 M5 Max MacBook Pro 上,Deltafin 在稳定解码时约能达到 1.00 token per second 。性能高度依赖存储吞吐,数据表明增加驱动器收益递减,因为每层需要读取的 16 个 expert reads 中最慢的那一项决定了总体速度。
虽然项目在优化上已有显著进展,但在一些技术难题上仍需突破,尤其是初始响应速度。目前由于在 prefill 阶段对模型层进行了重复读取,一个 512-token 的 prompt 在生成首个 token 前会有超过六分钟的延迟。开发者认为这是存储与架构上的瓶颈,而非硬件本身的限制,计划在后续迭代中通过优化来缩短这一延迟。
项目设计上提供了灵活的部署方式,既支持 full 安装,也支持 stream 模式。 stream 模式让用户以更小的存储起步,并在使用过程中逐步在本地缓存 experts 。 Deltafin 还内置了一个与 OpenAI 兼容的服务器,便于与各种客户端集成,并强制执行严格、贪婪且可重现的生成流程,以保证与模型预期输出一致。
总之,Deltafin 致力于在家用基础设施上拓展可能性。通过示范如何在不牺牲模型完整性的前提下运行大规模模型,项目希望为更广泛的 local-AI 社区做出贡献。开发者强调其核心任务是探索,并记录每次实验与基准测试,以透明披露在可及的高端消费者系统上运行 frontier models 的成本与能力。
Deltafin is an experimental, open-source project designed to run the full, 2.8-trillion-parameter Kimi K3 model locally on consumer hardware. Unlike other approaches that prune or quantize model weights to achieve performance, Deltafin prioritizes maintaining the original quality of the MoE (Mixture of Experts) model. By leveraging a single native binary, it ensures that Kimi K3 remains the sole authority for every generated token, with small draft models acting only as guessers that require explicit verification by the main model.
The current implementation focuses on extreme hardware efficiency, specifically targeting Apple Silicon setups with 128 GB of RAM and multiple SSDs for streaming expert weights. Benchmarks indicate that, when using an M5 Max MacBook Pro, Deltafin achieves approximately 1.00 token per second for steady decoding. The performance is heavily reliant on storage throughput, with data showing that adding more drives yields diminishing returns, as the total speed is constrained by the slowest of the 16 expert reads required for each layer.
While the project has made significant strides in optimization, there are still technical hurdles to overcome, particularly regarding initial processing speeds. A 512-token prompt currently faces a delay of over six minutes before the first token is generated due to redundant re-reads of model layers during the prefill phase. The developers have identified this as a storage and architectural bottleneck rather than a fundamental limitation of the hardware, and they have planned optimizations to address this latency in future iterations.
The project is structured to offer flexibility in how users deploy the model, providing both a "full" installation and a "stream" mode. The streaming option allows users to start with a smaller storage footprint and build a local cache of experts over time as they interact with the model. Additionally, Deltafin includes an OpenAI-compatible server, which allows it to integrate with various clients while enforcing a strict, greedy, and reproducible generation process that mirrors the model's intended output.
Ultimately, Deltafin serves as a research effort to push the boundaries of what is possible on home infrastructure. By demonstrating that large-scale models can be executed without compromising their integrity, the project aims to contribute to the broader local-AI community. The developers emphasize that their primary mission is exploration, documenting every experiment and benchmark to provide transparency into the costs and capabilities of running frontier models on accessible, albeit high-end, consumer systems.
• 在 MacBook Pro 配合外接 Thunderbolt SSD 等消费级硬件上,实现了对 2.78T 参数模型 Kimi K3 的每秒 token 推理,证明了在本地运行前沿规模模型的可行性。
• 当前系统架构需要从磁盘流式读取专家权重,这在预填充阶段形成严重瓶颈——每层需读取约 9TB 数据,极大地占用内存带宽。
• 性能扩展高度依赖磁盘吞吐量。所谓"drive ladder"显示,随着系统中 SSD 数量增加,收益递减,说明顺序带宽与调度策略比单纯增加硬件更为关键。
• 虽然当前推理速度对交互式聊天而言过慢,但这种配置适用于异步、无人值守的任务——模型可在夜间批量处理数据,从而实现数据不出机的本地高参数模型运行。
• 本地 LLM 性能的未来可能依赖定制的 ASIC 架构,例如晶圆级芯片,将极高的内存带宽置于目前 GPU 常见的商品化离散 DRAM 之上。
• 关于模块化 LLM 的设想颇多:通过索引"粘性"专家或权重,让模型的大量参数只需有一小部分驻留在 VRAM 中,从而有效降低万亿参数模型的使用门槛。
• 社区内部存在明显分歧:一方面优先考虑即时的实际应用,另一方面则更看重技术探索本身以及"试试看能不能做到"的精神。
• 在预填充调度方面做出重大优化(将每层多次读取专家改为单次顺序通过)是让这些"缓慢"的本地模型可用于诸如单 token 分类等任务的主要途径。
• 随着 AI 模型不断增强,编码与测试的工作流正在向人机协作转变,打字和调试的"成本"正在被迭代提示的成本所取代。
• 尽管对在本地运行模型存在普遍怀疑,但出于隐私保护、避免经常性订阅开支以及在本地硬件上利用大规模模型的需求,相关技术实验仍然受到强烈驱动。
讨论的核心在于:在消费级硬件上尝试运行万亿级参数模型的技术胆识,与当前速度限制之间的矛盾。参与者围绕这种缓慢推理的实际效用展开辩论,最终在强调本地数据主权而非实时响应的无人值守批处理工作流中找到了价值。这项努力既被视为对即时生产力的探索,也被看作迈向模块化、本地化 AI 的重要一步,呼应了计算史上"过去认为不可能的事情最终成为常态"的发展轨迹。
• Achieving token-per-second inference on a 2.78T parameter model (Kimi K3) using consumer hardware like a MacBook Pro and external Thunderbolt SSDs demonstrates the feasibility of running frontier-scale models locally.
• The current system architecture involves streaming expert weights from disk, which creates significant bottlenecks during the prefill phase where memory bandwidth is consumed reading 9TB of data per layer.
• Performance scaling is heavily dependent on disk throughput, with a "drive ladder" showing diminishing returns as more SSDs are added to the system, proving that sequential bandwidth and scheduling are more critical than raw hardware count.
• While the current inference speed is too slow for interactive chat, the setup is viable for asynchronous, unattended tasks where a model processes data overnight, allowing for local execution of high-parameter models without data leaving the machine.
• The future of local LLM performance likely involves custom ASIC architectures, such as wafer-scale chips, which prioritize extremely high memory bandwidth over the commodity discrete DRAM found in current GPUs.
• Speculation exists regarding modular LLM design, where "sticky" experts or weights could be indexed to allow only a fraction of a massive model to reside in VRAM, effectively democratizing the use of trillion-parameter models.
• There is a clear divide in the community between those who prioritize immediate practical application and those who value the spirit of technical exploration and "just seeing if it can be done."
• Significant optimizations in prefill scheduling—moving from reading experts multiple times per layer to a single-pass approach—represent the primary path toward making these "slow" local models useful for tasks like single-token classification.
• Coding and testing workflows are shifting as AI models become more capable, with a trajectory leading toward human-AI collaboration where the "cost" of typing and debugging is replaced by the cost of iterative prompting.
• Despite common skepticism about running models locally, the ability to maintain privacy, avoid recurring subscription costs, and utilize massive models on local hardware remains a compelling motivation for technical experimentation.
The discussion centers on the tension between the technical audacity of running trillion-parameter models on consumer hardware and the practical limitations of current speed benchmarks. Participants debate the utility of such slow inference, ultimately finding value in unattended, batch-style workflows that emphasize local data sovereignty over real-time responsiveness. This effort is framed not merely as a quest for immediate productivity, but as a fundamental step toward modular, local AI, mirroring the historical trajectory of computing where previously "impossible" tasks eventually became standard.
Meta 推出 Muse,这是一款用于管理日常事务并自动化处理各种琐事的个人 AI 代理。与以对话为主的传统 AI 助手不同,Muse 以代理式 AI 运行,能够执行多步任务、浏览网页并与多种第三方应用交互。它运行在一个配备专用浏览器的持久化虚拟机上,即便用户不在应用内,也能完成预约、填写在线表单或处理客服咨询等操作。 Meta has introduced Muse, a personal AI agent designed to manage daily tasks and automate busywork across various facets of life. Unlike traditional AI assistants that primarily handle conversational exchanges, Muse operates as an agentic AI, meaning it has the capability to perform multi-step tasks, navigate the web, and interact with various third-party applications. It functions on a persistent, dedicated virtual machine equipped with a specialized browser, allowing it to complete activities such as booking appointments, filling out online forms, or managing customer service inquiries even while the user is away from the app.
Meta 推出 Muse,这是一款用于管理日常事务并自动化处理各种琐事的个人 AI 代理。与以对话为主的传统 AI 助手不同,Muse 以代理式 AI 运行,能够执行多步任务、浏览网页并与多种第三方应用交互。它运行在一个配备专用浏览器的持久化虚拟机上,即便用户不在应用内,也能完成预约、填写在线表单或处理客服咨询等操作。
使用体验以简洁的对话交互为核心。用户可通过专用应用或 WhatsApp 用自然语言下达指令。收到目标或任务后,Muse 会制定行动计划、跟踪进度并主动推进必要步骤。如果完成某项任务所需的工具不存在,系统可以构建自己的工具,同时还能与 Email 、 Calendars 和 Instagram 等现有服务集成,保持工作流程的连贯性。
在安全和隐私方面,系统内置多重机制予以保障。 Muse 把用户登录凭据存放在安全的凭证库中,代理本身无法直接访问;并正在接入 1Password 等工具以进一步保护敏感信息。用户购物时,系统会生成一次性卡号,确保商家无法获取真实支付信息。符合条件的购买还受 Link 的购买保护覆盖;公司强调用户与 Muse 的对话不会用于 Meta 的广告系统。
控制权始终掌握在用户手中,系统在执行关键操作前会征求用户同意。无论是发送 Email 还是完成交易,用户都可以审阅代理拟定的计划、批准或拒绝操作,并查看代理的完整历史记录与即将执行任务的审计日志。权限可按连接逐一管理,便于用户在不同场景下灵活设定 AI 的自主程度。
Muse 目前提供限额的免费使用,并为需要更高使用量的用户提供付费订阅。作为持续运行的后台代理,它旨在将模式从被动辅助转变为主动管理——根据用户的具体目标与常设指令,持续监控信息并自动执行任务。
Meta has introduced Muse, a personal AI agent designed to manage daily tasks and automate busywork across various facets of life. Unlike traditional AI assistants that primarily handle conversational exchanges, Muse operates as an agentic AI, meaning it has the capability to perform multi-step tasks, navigate the web, and interact with various third-party applications. It functions on a persistent, dedicated virtual machine equipped with a specialized browser, allowing it to complete activities such as booking appointments, filling out online forms, or managing customer service inquiries even while the user is away from the app.
The user experience is centered around simple, conversational interactions. Users can communicate with Muse through a dedicated app or via WhatsApp, providing instructions in plain language. Once a goal or task is assigned, Muse formulates an action plan, tracks progress, and proactively manages the necessary steps. If a required tool for a specific task is unavailable, the system is designed to build its own tools, and it can integrate with existing services like email, calendars, and Instagram to maintain seamless workflows.
Security and privacy are prioritized through several built-in mechanisms. Muse stores user logins in a secure credential store that the agent itself cannot access, and it is adding integration for tools like 1Password to further protect sensitive information. When a user makes a purchase, the system generates a one-time card number to ensure that the merchant never gains access to the user's real payment details. Furthermore, eligible purchases are covered by Link's purchase protections, and the company emphasizes that user conversations with Muse are not shared with Meta's advertising systems.
Control remains firmly in the hands of the user, as the system is designed to seek approval before performing critical actions. Whether the task involves sending an email or completing a transaction, users can review the agent's proposed plan, approve or deny actions, and monitor a complete audit trail of the agent's history and upcoming tasks. These permissions can be managed on a per-connection basis, providing flexibility to dictate how much autonomy the AI has in different contexts.
Muse is currently available with a usage limit for free, with options for a paid subscription for users who require greater capacity. By functioning as a continuous background agent, it aims to shift the paradigm from simple assistance to active management, where the AI proactively monitors information and executes tasks based on the user's specific goals and standing instructions.
• Meta 的新 AI agent Muse 目标是占领主流"普通用户"市场,把自己定位为一个实用且易用的数字助理,作为出厂预装的工具而非小众技术产品。
• 技术圈的狂热与大多数普通用户之间存在巨大脱节:前者沉迷于追踪模型层级和参数,后者几乎不关心底层技术、公司品牌或所用 AI 模型的细微差别。
• 鉴于 Meta 在隐私丑闻方面的历史及其通过用户行为变现的商业模式,很多用户明显不信任 Meta 在个人数据管理上的做法。
• AI agent 的产品设计常陷入"高管泡沫",开发者往往把重点放在安排会议或预订行程等企业级任务上,这些本质上是高层管理人员的行政事务,而非普通消费者最迫切的需求。
• Agent 的有效性依赖于与个人数据(电子邮件、日历、购买记录)深度集成,这创造了一个高风险环境,容易发生提示注入、隐私泄露和对抗性操控等问题。
• 批评者认为消费类技术在很大程度上已经"基本解决",像 Meta 这样的公司在没有实际需求的地方强行加入 AI 以追求更多利润;支持者则强调 AI 在自动化处理当前工具难以协调的复杂多步物流任务时,确实具有实际价值。
• 对"个人助理"这一承诺的怀疑不断出现,很多人回想起过去失败的尝试(如 Facebook M),并指出这些工具可能更多地成为复杂的广告投放界面,而非真正有用的 agent 。
• 虽然技术精通的用户可能为了最大限度地控制和保护隐私而倾向于自托管或手动工作流,但在全球超过 30 亿的用户群体中,绝大多数人更看重便利性和省力,而非理论上的隐私担忧。
• Agentic AI 的快速商品化意味着即便 Meta 未能主导市场,整个生态也正向能够浏览网页、填写表格并执行任务的产品转移,使得手动浏览和人工交互在许多常见工作流中正变得愈发过时。
• 归根结底,此类产品的成功可能并不完全取决于技术优劣,而取决于 Meta 独特的杠杆力量——利用其庞大且根深蒂固的用户基础,在人们日常使用的平台上推广一个"开箱即用"的 agent 。
这场讨论反映出在 agentic AI 的巨大潜力与构建这些系统的公司间普遍存在的不信任之间,存在根深蒂固的张力。虽然技术型用户更看重控制、隐私和开放系统,但现实是大众在意识形态或安全顾虑面前,通常更倾向于易用性和集成化服务。最终,这些 AI agents 标志着互联网体验从手动、人为导航的工作流向自动化、目标导向交互的转变,即便像 Meta 这样的公司主要动机仍是扩展其广告和数据采集生态系统。
• Meta's new AI agent, Muse, aims to capture the mainstream, "normie" market by positioning itself as a helpful, easy-to-use digital assistant that functions as a factory-installed utility rather than a niche tech tool.
• A significant disconnect exists between the tech-enthusiast community, which obsessively tracks model tiers and parameters, and the vast majority of users, who are largely oblivious to the underlying technology, company brands, or the nuances of AI model selection.
• Many users perceive a glaring lack of trust regarding Meta's stewardship of personal data, given the company's history of privacy scandals and its business model of monetizing user behavior.
• Product design for AI agents often suffers from "executive bubble" syndrome, where developers focus on corporate tasks like scheduling meetings or booking travel, which are essentially the mundane administrative duties of senior management rather than the most pressing needs of everyday consumers.
• The effectiveness of an agent relies on deep integration with personal data (email, calendar, purchase history), creating a "hazmat" environment where the risk of prompt injection, privacy leaks, and adversarial manipulation is high.
• Critics argue that consumer technology is already largely "solved," and that firms like Meta are forcing AI integration into non-problems to extract more profit, while supporters highlight genuine utility in automating complex, multi-step logistical tasks that current tools struggle to coordinate.
• There is a recurring skepticism toward the "personal assistant" promise, with many recalling previous failed attempts (like Facebook M) and noting that these tools may function primarily as sophisticated ad-targeting surfaces rather than helpful agents.
• While the tech-savvy segment of the population may prefer self-hosted or manual workflows for maximum control and privacy, the vast majority of the 3+ billion-strong global user base likely prioritizes convenience and energy minimization over theoretical privacy concerns.
• The rapid commoditization of agentic AI means that even if Meta fails to capture the market, the landscape is shifting toward products that can navigate the web, fill forms, and execute tasks, making manual browsing and interaction increasingly obsolete for many common workflows.
• Ultimately, the success of such products may hinge not on technical superiority, but on Meta's unique ability to leverage its massive, entrenched user base to distribute an agent that "just works" within the platforms people already use daily.
The discussion reflects a deep-seated tension between the immense utility of agentic AI and the pervasive distrust of the corporations building it. While technical users prioritize control, privacy, and open systems, they often fail to account for the reality that the general public consistently favors ease of use and integrated services over ideological or security-based objections. Ultimately, these AI agents represent a shift in the internet experience from manual, human-navigated workflows to automated, goal-oriented interaction, even if the primary motivation for companies like Meta remains the expansion of their advertising and data-harvesting ecosystems.
OpenAI 已推出 ChatGPT Images 2.5,一款旨在提升数百万用户创意工作流程的先进模型。该版本在细节清晰度、纹理丰富度和光影自然度上都有显著提升。值得注意的是,与上一代相比,生成延迟最高降低了 50%,让用户能更快地迭代视觉创意。 OpenAI has introduced ChatGPT Images 2.5, a state-of-the-art model designed to improve the creative workflow for millions of users. This iteration focuses on generating images with sharper details, richer textures, and more natural lighting. Notably, it also achieves up to a 50% reduction in generation latency compared to its predecessor, allowing users to iterate on their visual concepts much faster.
OpenAI 已推出 ChatGPT Images 2.5,一款旨在提升数百万用户创意工作流程的先进模型。该版本在细节清晰度、纹理丰富度和光影自然度上都有显著提升。值得注意的是,与上一代相比,生成延迟最高降低了 50%,让用户能更快地迭代视觉创意。
该模型在使用参考照片时更能保持主体一致性,并能在多轮对话中更可靠地执行编辑指令。为增强 ChatGPT 内的创作体验,公司新增了如 Sketch 的工具,允许用户在对话中直接绘制,作为最终输出的视觉参考。其它功能还包括针对传单等常用格式的模板、可在图像上直接添加注释以便集中编辑,以及新的共享选项,允许用户附上原始提示词,便于他人基于其继续创作。
面向开发者,OpenAI 通过 API 发布了两款不同的模型:GPT-Image-2.5 Flare 作为多数应用的标准高性能选择,兼顾速度与质量;GPT-Image-2.5 Sunburst 则面向需要高精度编辑与控制的专业创意工作流程。这些工具已被 Adobe 、 Runway 和 Higgsfield AI 等公司整合,以简化生产与专业创作任务。
安全仍是此次发布的核心,新模型集成了现有的安全防护、提示词与图像检测机制,并支持 C2PA 元数据以确保透明度。该更新已向所有 ChatGPT 层级的用户开放(包括桌面端和移动端),开发者也可通过 API 平台立即访问这些新模型。
OpenAI has introduced ChatGPT Images 2.5, a state-of-the-art model designed to improve the creative workflow for millions of users. This iteration focuses on generating images with sharper details, richer textures, and more natural lighting. Notably, it also achieves up to a 50% reduction in generation latency compared to its predecessor, allowing users to iterate on their visual concepts much faster.
The model is built to better maintain subject consistency when working from reference photos and provides more reliable adherence to editing instructions over multi-turn conversations. To enhance the creative process within ChatGPT, the company has added new tools like Sketch, which allows users to draw directly in the chat as a visual guide for the final output. Other features include templates for popular formats like flyers, the ability to add direct comments to images for focused editing, and new sharing options that let users include their original prompts so others can build upon them.
For the developer community, OpenAI is releasing two distinct models via the API. GPT-Image-2.5 Flare serves as the standard, high-performance choice for most applications, offering improved speed and quality. Meanwhile, GPT-Image-2.5 Sunburst is tailored for premium creative workflows that require high-precision editing and control. These tools are already being integrated by companies such as Adobe, Runway, and Higgsfield AI to streamline production and professional creative tasks.
Safety remains a core component of this release, with the new model incorporating existing safeguards, prompt and image checks, and C2PA metadata to ensure transparency. The update is currently available to users across all ChatGPT tiers, including desktop and mobile, while developers can access the new models immediately through the API platform.
图像生成技术被广泛用于日常个人事务,例如设想家居翻新、园艺布置,以及为家人和朋友制作轻松的内容。
公众对这项技术的看法严重两极分化:一方面有人为其带来的新型创作能力和"神奇"的生产力工具感到兴奋;另一方面则对大量"粗制滥造"("slop")、错误信息的激增以及大规模数据中心对环境的影响深表担忧。
一个重要争议点在于 AI 生成内容在商业场景中的常态化。许多人对餐单和本地广告中使用低质量或具有误导性的 AI 图像感到沮丧和反感。
关于资源消耗的争论尚未平息:有人认为数据中心的能耗与高尔夫球场维护或个人出行等其他社会习惯相比微不足道;也有人主张所有数字内容都应强制披露碳排放信息,要求透明化。
高质量图像处理的普及挑战了传统的"真实性"与证据观念:照片被有效地变成了可塑的数字文件,不再能作为现实的不可变证明。
一些用户发现将 AI 作为认知或创意辅助工具极有价值,尤其对那些患有意象缺失症(aphantasia)的人,或对那些希望在不受高门槛技术限制下快速原型化视觉构想的人来说,帮助尤为明显。
批评者则强调,图像生成的高速与大量产出助长了无脑消费与"劣质创作"的文化,认为创作的简易性削弱了人类艺术表达和真实体验的价值。
对现有模型技术局限性的质疑仍然存在。人们反复发现诸如解剖细节错误(例如手指数目不对)等顽固缺陷,且模型往往倾向生成一种独特而重复的审美风格。
隐私与同意问题也极为突出,尤其是在用 AI 工具改变或"混合再创作"儿童与家人照片时,这引发了关于长期数据安全与数字纪念伦理的广泛担忧。
当人们将 AI 的影响与其他个人或企业习惯相比较时,关于"whataboutism"(转移论证)与伪善的争论经常出现,这反映出一个根本分歧:当前的生态伤害是否足以证明引入新技术是合理的。
总体来看,这场讨论反映出两股尖锐对立的声音:一方面有人把生成式 AI 视为赋能性的创意突破,另一方面有人将其视为对视觉完整性与环境稳定性的存在性威胁。虽然许多人已将这些工具整合进日常生活,用于家居装饰和个人叙事等看似无害且实用的用途,但面对社交媒体和商业平台上合成内容的泛滥,公众明显感到疲惫。讨论的走向取决于相互冲突的价值观——对个体创作民主化的渴望与对一个充斥着难以分辨且易产生错误信息图像的社会的恐惧相互对立。话语中的模式显示,尽管开发者欢迎速度与分辨率等技术改进,但该技术在更广泛文化层面的同化仍极具争议,且容易引发反复的伦理争论。
• Image generation technology is widely used for mundane, personal tasks such as visualizing home renovations, gardening layouts, and creating lighthearted content for family and friends.
• Perspectives on the technology are deeply polarized, ranging from enthusiasm for new creative capabilities and "magical" productivity tools to deep concern over the proliferation of "slop," misinformation, and the environmental impact of large-scale data centers.
• A significant point of contention involves the normalization of AI-generated content in commercial spaces, with many expressing frustration over the use of low-quality or misleading AI imagery in restaurant menus and local advertisements.
• The debate surrounding resource consumption remains unresolved, with some arguing that data center energy use is trivial compared to other societal habits like golf course maintenance or individual travel, while others advocate for mandatory carbon transparency for all digital content.
• The democratization of high-quality image manipulation creates challenges for traditional notions of truth and evidence, effectively turning photographs into malleable digital files that no longer serve as immutable proof of reality.
• Some users find immense value in using AI as a cognitive or creative aid, particularly those with aphantasia or those seeking to prototype visual ideas without the barrier of high-level technical skill.
• Critics emphasize that the speed and volume of image generation encourage a culture of mindless consumption and "slop," arguing that the ease of creation degrades the value of human artistic expression and authentic experiences.
• Skepticism persists regarding the technical limitations of current models, with recurring observations about lingering flaws such as inaccurate anatomical details (e.g., finger counts) and the tendency for models to produce a distinct, repetitive aesthetic.
• Concerns over privacy and consent are prominent, particularly regarding the use of AI tools to alter or "remix" photos of children and family members, raising questions about long-term data security and the ethics of digital memorialization.
• Arguments regarding "whataboutism" and hypocrisy frequently emerge when participants compare the impact of AI to other personal or corporate habits, highlighting a fundamental disagreement over whether current ecological harm justifies the introduction of new technologies.
The conversation reflects a sharp divide between those who view generative AI as an empowering, creative breakthrough and those who perceive it as an existential threat to visual integrity and environmental stability. While many individuals integrate these tools into their daily lives for harmless, practical purposes like home decoration and personal storytelling, there is a palpable fatigue regarding the flood of synthetic content across social media and commercial platforms. The discussion ultimately hinges on conflicting values, pitting the desire for individual creative democratization against the fear of a society saturated by indistinguishable, misinformation-prone imagery. Patterns in the discourse suggest that while technical improvements like speed and resolution are welcomed by builders, the broader cultural assimilation of the technology remains highly contentious and prone to recurring ethical debates.
OpenAI 宣布在数学领域取得重大突破,解决了 Navier.Stokes existence and smoothness problem,这是七个 Millennium Prize Problems 之一。近约 90 年来,数学家们一直争论三维平滑流体运动是否会在有限时间内崩解为奇异性,即流体速度不受约束地趋于无限。研究人员借助一款先进且尚未公开的内部 AI 模型,给出了一个解析证明,并在 Lean 中完成了形式化证明,表明在平滑力的作用下,最初平滑的流体确实可能出现这种奇异性。 OpenAI has announced a significant breakthrough in mathematics by resolving the Navier.Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. For approximately 90 years, mathematicians have debated whether smooth three-dimensional fluid motion can break down into a singularity, where fluid speeds grow without bound in finite time. By utilizing an advanced, unreleased internal AI model, researchers produced both an analytical proof and a formalization in Lean, demonstrating that such a singularity can indeed occur in an initially smooth fluid under the influence of a smooth force.
OpenAI 宣布在数学领域取得重大突破,解决了 Navier.Stokes existence and smoothness problem,这是七个 Millennium Prize Problems 之一。近约 90 年来,数学家们一直争论三维平滑流体运动是否会在有限时间内崩解为奇异性,即流体速度不受约束地趋于无限。研究人员借助一款先进且尚未公开的内部 AI 模型,给出了一个解析证明,并在 Lean 中完成了形式化证明,表明在平滑力的作用下,最初平滑的流体确实可能出现这种奇异性。
证明的核心在于一个涡旋的演化:涡旋向内螺旋并被拉长,使流体速度在能量保持有限的情况下趋于无限。技术上的关键是 Navier.Stokes equations 中力的精确平衡——加速度、压力梯度与粘性之间的相互作用共同促成了这一种解的破裂。由此该研究表明,流体运动的连续体近似最终可能失效,后续演化必须转向分子层面的建模才能描述。
为实现这一成果,OpenAI 使用了由其最新高性能内部模型驱动的多智能体系统。约 10,000 个并发智能体组成的大型群体,在代码执行与互联网检索等工具的辅助下,经过数日协作,反复迭代各种问题表述。该过程始于对 Euler regularity problem 的一次成功尝试,这成为通向 Navier.Stokes equations 解决方案的垫脚石。智能体被鼓励探索多条路径,并定期用 Codex 汇总见解以打磨最终证明。
组织澄清称,他们参与该项目是被新一代 AI 模型所展现的前所未有能力所驱动,而非为争取 Millennium Prize 。尽管另有独立研究小组也在攻关相关问题,OpenAI 表示他们的发现是通过多智能体系统与人类研究者的独立努力取得的。在 Lean 中对证明的形式化验证凸显了 AI 作为严谨数学研究伙伴的日益价值。
这一里程碑反映了人工智能领域正在发生的快速变革。 OpenAI 强调,这不是终点,而是一个需要谨慎引导的新进步时代的标志。随着能力的持续推进,他们将致力于模型的可控、负责任且具备问责机制的发展,以确保这些进步惠及全人类。
OpenAI has announced a significant breakthrough in mathematics by resolving the Navier.Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. For approximately 90 years, mathematicians have debated whether smooth three-dimensional fluid motion can break down into a singularity, where fluid speeds grow without bound in finite time. By utilizing an advanced, unreleased internal AI model, researchers produced both an analytical proof and a formalization in Lean, demonstrating that such a singularity can indeed occur in an initially smooth fluid under the influence of a smooth force.
The proof centers on the behavior of a vortex, which spirals inward and elongates in a way that allows the fluid's velocity to reach infinite speeds while maintaining finite energy. The technical achievement lies in the precise balancing of forces within the Navier.Stokes equations, where acceleration, pressure gradients, and viscosity interact to permit this breakdown. By establishing this outcome, the research confirms that the continuum approximation of fluid motion can eventually fail, requiring a shift to molecular-level modeling to track the system's further evolution.
To achieve this, OpenAI employed a multiagent system powered by their latest, highly capable internal model. A large group of approximately 10,000 concurrent agents, assisted by tools like code execution and internet search, collaborated over several days to iterate on various problem formulations. The process began with a successful attempt at the Euler regularity problem, which served as a stepping stone that eventually led the agents to the solution for the Navier.Stokes equations. The agents were encouraged to explore diverse approaches, with insights periodically consolidated using Codex to refine the final proof.
The organization clarified that their interest in this project was driven by the unprecedented performance of their new AI models rather than a desire to claim the Millennium Prize itself. While a separate team of researchers had been working on a related problem, OpenAI confirmed that their own discovery was reached independently through the efforts of their multiagent system and human researchers. The successful verification of the proof via Lean highlights the increasing utility of AI as a partner in rigorous mathematical discovery.
Ultimately, this milestone serves as a snapshot of the rapid evolution currently unfolding in the field of artificial intelligence. OpenAI emphasizes that this is not a final destination, but rather evidence of a new era of progress that requires careful stewardship. As the organization continues to advance its capabilities, it intends to focus on the responsible, steerable, and accountable development of these models to ensure their growth benefits humanity at large.
• 数学研究员 Tristan Buckmaster 和 Levent Alpöge 指控 OpenAI 施压,要求将 Levent Alpöge 从作者名单中移除,并推动以 OpenAI 为主导的叙事,事由该模型声称已给出 Navier-Stokes 问题的解答。研究人员认为这是 OpenAI 试图在受其自身工作影响的研究成果中争取优先权。
• OpenAI 承认其服务中去标识化的用户数据可能被用于模型训练,但坚称其证明是独立完成的,并且在科学层面上与研究人员关于 Euler 问题的工作截然不同。
• 利用内部模型和 10,000 个并发代理快速攻克 Millennium Prize 问题,凸显了 AI 研究能力的巨大跃升,也引发了激烈争论:这到底预示着通用人工智能的到来,还是只是算力驱动的蛮力式突破。
• 怀疑者认为,关于训练数据被污染"可能性不大"的说法是含糊其辞的借口,意在掩盖潜在的知识产权挪用。他们指出,AI 公司可以通过监控专有资料并扩展算力,在原始创新者正式公开成果之前"抢跑",从而在实质上超越研究人员。
• 对"认知黑暗森林"式风险的担忧与日俱增:在这种环境下,AI 驱动的平台可能窃取或抢先发表独特见解,迫使研究人员和企业更倾向于极端保密而非公开协作。
• 从第一天起就用 Lean proof assistant 对证明进行机器验证被视为积极进展,为验证数学结论提供了客观的、可检验的基准。
• 有观察者指出,人们对优先权纠纷的关注掩盖了一个更为根本且足以改变世界的事实:AI 模型已经达到或超过了此前预期数年后才能见到的数学能力水平。
• 关于 OpenAI 企业文化是否本质上存在有害倾向的伦理担忧依然存在,尤其是针对其高压策略以及对学术界在署名和协作规范方面表现出漠视的报道。
• 此事件提出了关于数据主权的关键问题。许多人主张,研究人员和企业应停止将敏感或未发表工作交由基于云的 AI 工具处理,因为现行服务条款通常允许提供商利用用户数据进行训练并可能复制用户见解。
• 许多参与者认为,这可能是知识产权保护的转折点,AI 自动剽窃时代或将迫切需要新的可追溯性工具,例如对训练数据进行滚动哈希,以证明由模型生成突破性成果的来源。
关于 Navier-Stokes existence and smoothness problem 解决方案的披露,成为人工智能史上的分水岭,前所未有的科学进展前景与现代研究竞争的现实发生了冲突。尽管这一技术成就被视为划时代的突破,但围绕抢发成果、强制删除署名以及潜在训练数据污染的争议,已严重削弱了 AI 开发者与学术界之间的信任。此次事件反映出一种日益加深的焦虑:随着这些系统愈发强大,它们可能从协作工具转变为竞争威胁,通过利用用户提供的数据来"抢跑"人类创新。归根结底,这一事件强调了一个越来越清晰的共识:未来的知识工作必须对数据隐私进行彻底重审,因为大规模模型聚合私人见解并据此行动的能力,使得传统的"开放"研究模式变得愈发岌岌可危。
• Mathematical researchers Tristan Buckmaster and Levent Alpöge allege that OpenAI attempted to pressure them into excluding Alpöge from authorship and adopting an OpenAI-led narrative following the model's Navier-Stokes solution, an action the researchers view as an attempt to assert priority over work influenced by their own efforts.
• OpenAI acknowledges the possibility that de-identified user data from its services may have informed its model's training, though it maintains that its proof is independent and scientifically distinct from the researchers' work on the Euler problem.
• The rapid resolution of a major Millennium Prize problem via an internal model and 10,000 concurrent agents highlights a massive leap in AI research capabilities, prompting significant debate over whether this represents the arrival of AGI or merely a brute-force application of compute.
• Skeptics argue that the "unlikely" claim regarding training data contamination is a "weasel-worded" admission of potential intellectual property appropriation, suggesting that AI companies can effectively "front-run" researchers by monitoring proprietary data and scaling compute to solve problems before the original innovators.
• Concerns are mounting regarding the "cognitive dark forest," where the risk of AI-powered platforms stealing or scooping unique insights incentivizes researchers and businesses to operate in extreme secrecy rather than collaborating openly.
• The use of the Lean proof assistant to verify the proof from day one is seen as a positive development, providing an objective, machine-checked baseline for verifying the validity of the mathematical result.
• Some observers suggest that the focus on the priority dispute masks the more fundamental, world-altering fact that an AI model has achieved a level of mathematical capability that few expected to see for several more years.
• Ethical concerns remain regarding whether OpenAI's corporate culture is fundamentally toxic, specifically citing reports of high-pressure tactics and dismissive attitudes toward the academic community's norms regarding research credit and collaboration.
• The incident raises critical questions about data sovereignty, with many arguing that researchers and enterprises should stop using cloud-based AI tools for sensitive or pre-publication work, as current terms of service often allow the providers to train on and potentially replicate user insights.
• Many participants view this as a potential tipping point for intellectual property, suggesting that the era of AI "slop" or automated plagiarism may necessitate new tools for traceability, such as rolling hashes of training data, to prove the provenance of generated breakthroughs.
The disclosure of a solution to the Navier-Stokes existence and smoothness problem marks a watershed moment in the history of artificial intelligence, forcing a collision between the promise of unprecedented scientific advancement and the realities of modern research competition. While the technical accomplishment is being hailed as an epochal achievement, the surrounding controversy over allegations of scooping, forced credit removal, and potential training data contamination has significantly damaged the trust between AI developers and the academic community. The discussion reflects a deepening anxiety that as these systems become more powerful, they may transition from collaborative tools to competitive threats that "front-run" human innovation by exploiting the very data users provide. Ultimately, the episode underscores a growing consensus that the future of knowledge work will require a radical re-evaluation of data privacy, as the ability of large-scale models to aggregate and act upon private insights renders the traditional "open" research model increasingly precarious.
108 comments • Comments Link
在量子实验的自动化校准方面(例如量子比特启用 qubit bring-up 和门操作表征 gate characterization),传统的脚本编写和数据记录就能高效完成,说明对这些常规任务来说,复杂的 AI 并非必需。
关于平台上伪造草根运动(astroturfing)和政治偏见的担忧,引发了社区文化演变的大讨论;一些用户认为它正从一个专业的科技社区,向更两极化、类似 Reddit 的环境转变。
对 AI 的怀疑常被解读为对美国企业横行霸道和权力集中化的理性批评,也有人把它归咎于个人的确认偏差——即用户感觉自己的观点受到更多敌视。
在开发工作流中采用 AI 的效果呈现两极化:有人报告功能交付上获得明显的生产力提升,另一些人则批评 AI 生成的输出冗长、难以理解,并且在维护时易出错。
目前 AI 领域资本的激增,尤其是在数据中心和专用硬件方面,被一些人视为类似以往的经济泡沫:大量资金流向未经验证、投资回报率(ROI)不明的未成熟技术。
关于大型语言模型(LLMs)的有效性也存在争议:它们既可以成为让专家在创作过程中保持控制权的强大工具,也可能被缺乏架构理解的人当作粗糙、低质量内容的来源。
当复杂议题被转为对公司宣传、上市时机或意识形态部落主义的元评论,而不是回到原始资料进行实质性讨论时,人们便担心批判性思维在衰退。
围绕"人类在环"(human-in-the-loop)模型,持续存在紧张关系:支持者称 AI 是对人类劳动力的倍增器,反对者则担心专业知识会被系统性替代,以及这种技术依赖对产业长期经济稳定性的影响。
把 AI 比喻成不可控的物理力量(比如滚落山坡的巨石),反映出一种日益增长的焦虑:技术威力巨大,却可能脱离人类监管,令人担忧。
即便是基于科学研究的关于 AI 的讨论,也越来越常被置于怀疑的视角:观察者往往将公开声明解读为影响市场估值的策略举措,而非纯粹的学术或技术进展。
这场讨论反映出社区内部关于 AI 发展方向和平台自我认同的严重分裂。部分人强调技术的实用性和生产力收益;另一些人则以深刻的犬儒主义质疑企业动机,将当下视为投机泡沫或智识话语的衰落。把 AI 看作是扩展人类能力的工具者,与把它视为危险且未知的破坏性力量者之间的紧张,折射出适应快速技术变迁时更广泛的文化冲突。最终,这场对话也表明这些专业性进展已被高度个人化和政治化,常常掩盖了被讨论主题的技术细节。 • Automating quantum experiment calibration, such as qubit bring-up and gate characterization, is highly effective using traditional scripting and data logging, suggesting that sophisticated AI is not strictly necessary for these routine tasks.
• The perception of astroturfing and political bias on the platform has fueled significant debate regarding the site's cultural evolution, with some users observing a transition from a specialized tech community to a more polarized, Reddit-like environment.
• Skepticism toward AI is often framed as either a rational critique of American corporate "steamrolling" and centralization, or a consequence of individual confirmation bias where users perceive more hostility toward the viewpoints they hold.
• AI adoption in development workflows is polarizing, with some reporting significant productivity gains in feature delivery, while others criticize the resulting output as verbose, unintelligible, and prone to breaking during maintenance.
• The current surge in AI capital investment, particularly in data centers and specialized hardware, is viewed by some as an economic bubble reminiscent of previous cycles, where massive capital is allocated to immature technologies with unproven ROI.
• The effectiveness of LLMs is debated as a spectrum, where they function as powerful tools for experts who maintain control over the creative process versus being used as a source of "slop" by those lacking deep architectural understanding.
• Concerns over the "death of critical thinking" are raised when complex scientific topics are diverted into meta-commentary about company propaganda, IPO timing, or ideological tribalism rather than engagement with the source material.
• There is a recurring tension regarding the "human-in-the-loop" model, where proponents argue for AI as a force multiplier for human labor, while detractors worry about the systematic displacement of professional expertise and the long-term economic stability of tech-dependent industries.
• Metaphors comparing AI to uncontrollable physical forces—like a boulder rolling down a mountain—highlight a growing anxiety that the technology is both significantly powerful and alarmingly detached from human oversight.
• Discussions about AI, even those grounded in scientific research, are increasingly viewed through a lens of suspicion, where observers interpret public announcements as strategic maneuvers designed to influence market valuation rather than purely academic or technical progress.
The discussion reflects a deep fragmentation within the community regarding the trajectory of AI development and the platform's own identity. While some participants emphasize technical utility and productivity gains, others express profound cynicism about the underlying corporate motivations, often framing the current era as a speculative bubble or a degradation of intellectual discourse. The tension between those who see AI as an essential human-extending tool and those who fear it as a disruptive, poorly understood force underscores a broader cultural struggle to adapt to rapid technological change. Ultimately, the conversation highlights how deeply personal and political these professional developments have become, often overshadowing the technical nuances of the subjects at hand.