I resigned from Anthropic today
734 points
• 5 days ago
• Article
Link
Jacob Coxon 已公开辞去在 Anthropic 的职位,结束了他在 Anthropic 和 OpenAI 从事预训练研究的三年任期。离职时,他对行业走向发出严正警告,认为这两家公司都未以负责任的方式行事。他指出,这些实验室正卷入争先研发可自我改进的超级智能的竞赛,实际上是在拿全球安全冒险一搏。
Coxon 担心,那些可能带来革命性影响的超人类系统正迅速发展,影响范围从高级黑客技术到夺取大量资源与权力。他强调这种进展并未放缓。尽管这些组织在外界看来波澜不惊,但他认为真正参与构建这些系统的人确实担心在十年内会发生灾难性后果,把这种威胁视为现实风险而非宣传噱头。
尽管内部存在如此严重的恐惧,开发仍在继续,两家主要实验室的原因各有不同。 Coxon 认为在 OpenAI,员工并未把这一文明风险真正内化;而在 Anthropic,风险虽然被充分认识,但公司陷入争先实现超级智能的竞赛,抱着"必须赢,因为不能指望别人能安全完成"的心态。
Coxon 抨击这种做法是傲慢的赌博,认为人类的未来不应由私人公司的内部决策来决定。他质疑所谓加速对齐的做法,指出那需要一种非凡的自信——相信不存在更好、更安全的路径。他呼吁美国各研究所加强协调,甚至提出可能需要更激进的措施,例如暂时中止开发,以阻止导致不可控系统的全球竞赛。
在结尾,他敦促同行研究者重新审视参与这些项目的决定,质问在没有对底层系统有基本而严谨理解的情况下就启动超级智能强化学习实验是否负责任。他劝告同行不要抱有发展不可避免的宿命论,而应在这一关键时刻停下来,深思他们工作的严重后果。
Jacob Coxon has publicly resigned from his position at Anthropic, concluding a three-year tenure spent conducting pretraining research at both Anthropic and OpenAI. In his departure, he issued a stark warning regarding the industry's direction, arguing that neither company is behaving in a responsible manner. He contends that these labs are currently engaged in a high-stakes race toward self-improving superintelligence, effectively gambling with global safety in the process.
The core of Coxon's concern lies in the rapid development of superhuman systems capable of revolutionary impacts, from advanced hacking to the acquisition of significant resources and power. He insists that this progress is not decelerating. While many within these organizations project a sense of outward composure, he maintains that the individuals responsible for building these systems genuinely fear the potential for catastrophic outcomes before the end of the decade, viewing the threat as a reality rather than a marketing tactic.
The reasoning behind why this development continues despite such severe internal fears differs between the two major labs. At OpenAI, Coxon suggests a lack of deep internalization regarding the civilizational stakes. Conversely, at Anthropic, the risks are well-understood, but the company is locked in a competitive race to reach superintelligence first, operating under the belief that they must win because they cannot rely on others to do so safely.
Coxon criticizes this approach as a form of hubristic gambling, asserting that the future of humanity should not be determined by the internal decision-making processes of a private company. He challenges the notion of a fast-tracked alignment process, suggesting that it demands an extraordinary level of confidence that no better, safer trajectories exist. He advocates for increased coordination between U.S. labs and raises the possibility that more drastic measures, such as a temporary moratorium on development, might be necessary to prevent a global race toward potentially uncontrollable systems.
Concluding his message, Coxon urges fellow researchers to reconsider their commitment to these projects. He questions whether it is responsible to initiate superintelligent reinforcement learning runs without a fundamental, rigorous understanding of the underlying systems. Instead of adopting a fatalistic attitude that these developments are inevitable, he encourages his peers to pause and reflect on the grave implications of their work during this critical period.
1014 comments • Comments Link
辩论的核心在于:AI 的发展究竟是否构成与核武器或气候变化相当的生存威胁,还是这些说法只是为了噱头或类宗教式的信号传递而夸大的"末日主义"。
怀疑派认为,大型语言模型(LLMs)不过是复杂的统计预测器,而非具有自主性的智能体。他们强调,目前的风险主要来自人类的滥用——而非机器自身产生了某种"意志"或感知能力。
持生存风险观点的人则认为,快速的规模化和自主"agentic"循环的出现——即模型能够与工具交互、执行代码并在多个系统中持续运行——开辟了一条通往灾难的路径。一旦达到某个临界点,这条路径可能变得难以预测、难以控制。
对于现有技术障碍(如样本效率限制、生物合成的难度和计算上的物理瓶颈)是否足以成为阻止 AI 驱动破坏的有效防线,双方存在严重分歧。
一个反复出现的主题是机构问责的问题。批评者指出,真正的危险在于权力集中于少数科技公司高管之手:这些人一方面鼓吹耸人听闻的场景,另一方面又在加速发展以谋取 IPO 回报或巩固市场主导地位。
关于劳动力市场的影响,观点也各不相同:有人认为 AI 是能提高人类生产力的有益工具;也有人担心它会通过高效地替代人类的认知劳动,造成一个被剥夺权利的永久性下层阶级。
这场对话反映出一种深刻矛盾:全球技术军备竞赛中那种"快速推进"的冲动,与现有治理和安全框架无法应对可能迅速显现后果之间存在根本冲突。
怀疑派指出,许多"末日"情景依赖科幻式比喻而非可观测证据。他们注意到,过去关于"有感知"模型的预测多次落空,这暗示当前的担忧可能同样缺乏依据。
一些参与者强调:如果灾难的概率确实很高,仅靠个别人员辞职或模糊的公开警告是远远不够的——这表明部分话语更多是在进行地位信号的传递,而非真正的安全策略。
讨论最终暴露出一个深刻的分歧:一方以传统工程和风险管理的视角看待 AI,认为它是软件工程可预见的延续,现有市场力量和物理约束可以防止失控;另一方则把 AI 视为一种变革性的相变,认为这种转变使得传统历史先例和安全指标失效。对他们而言,自主型 agent 框架与大规模计算的结合,会产生本质上难以侦测的风险,直到这些风险显现为生存威胁。
在这些技术争论之下,还潜藏着对行业领袖动机的深度怀疑。各方达成的一项广泛共识是:当前的治理机制既无法应对灾难性场景,也难以应对已经开始显现的重大社会动荡。 The central point of debate concerns whether AI development poses an existential threat comparable to nuclear weapons or climate change, or if such claims are hyperbolic "doomerism" used for marketing or religious-like signaling.
Skeptics argue that LLMs are merely sophisticated statistical predictors rather than autonomous agents, emphasizing that risks are currently limited to human misuse rather than emergent machine "will" or sentience.
Proponents of the existential risk perspective contend that rapid scaling and the emergence of autonomous "agentic" loops—where models can interact with tools, execute code, and persist across systems—create a path to catastrophe that is impossible to predict or contain once it reaches a certain threshold.
Significant disagreement exists regarding whether current technological hurdles, such as sample efficiency, the difficulty of biological synthesis, and the physical limitations of compute, serve as meaningful barriers to AI-driven destruction.
A recurring theme is the role of institutional accountability. Critics suggest that the real danger lies in the concentration of power among a small group of tech executives who are simultaneously promoting alarmist scenarios while accelerating development to secure IPO gains or market dominance.
Perspectives on labor market impact range from the view that AI acts as a beneficial tool for human productivity to fears that it will create a permanent, disenfranchised underclass by effectively replacing the value of human intellectual labor.
The conversation reflects a deep tension between the urge to "move fast" in a global technological arms race and the inability of existing governance and safety frameworks to manage outcomes that could manifest rapidly.
Skeptics point out that many "doomer" scenarios rely on science-fiction tropes rather than observable evidence, noting that past predictions of "sentient" models have consistently failed to materialize, suggesting that current concerns may be similarly unfounded.
Some participants highlight that if the potential for catastrophe is genuinely high, the current reliance on individual resignations or vague public warnings is an insufficient response, suggesting that the discourse is more about status signaling than practical safety.
The discussion ultimately reveals a profound divide: one group views AI through the lens of traditional engineering and risk management, while the other views it as a transformative phase change that renders traditional historical precedents and safety metrics obsolete.
The discussion highlights a fundamental polarization in how society perceives the trajectory of artificial intelligence. One camp views current progress as a predictable continuation of software engineering, where existing market forces and physical constraints prevent runaway outcomes. The opposing side treats AI as a potential "black swan" event, arguing that the combination of autonomous agentic frameworks and massive computational scaling creates risks that are inherently undetectable until they manifest as existential threats. Beneath these technical arguments lies a deep skepticism regarding the motivations of the industry leaders, with a broad consensus that current governance is ill-equipped to address either the catastrophic scenarios or the significant societal disruptions already beginning to unfold.