尽管 Large Language Models 在生成语法正确的代码方面已非常娴熟,但它们的普及也带来了一个隐性的代码质量危机。即便代码能通过自动化测试,往往仍存在过度抽象、冗余重复和糟糕的架构选择。此类问题通常被称为 slop,会导致代码行数失控式增长。随着项目每月新增数百万行,人类开发者已难以保持监督,而与常见观点相反,AI agents 目前也无法自我修正这种技术债务的累积。 While Large Language Models have become remarkably adept at generating formally correct code, their rise has ushered in a hidden crisis of code quality. Even when code passes automated tests, it often suffers from excessive abstraction, redundant duplication, and poor architectural decisions. This phenomenon, often termed slop, leads to an uncontrolled explosion of lines of code. As projects grow by millions of lines per month, human developers lose the ability to maintain oversight, and contrary to popular belief, AI agents are currently incapable of self-correcting this accumulation of technical debt.
尽管 Large Language Models 在生成语法正确的代码方面已非常娴熟,但它们的普及也带来了一个隐性的代码质量危机。即便代码能通过自动化测试,往往仍存在过度抽象、冗余重复和糟糕的架构选择。此类问题通常被称为 slop,会导致代码行数失控式增长。随着项目每月新增数百万行,人类开发者已难以保持监督,而与常见观点相反,AI agents 目前也无法自我修正这种技术债务的累积。
目前业界对这一问题的评估过于依赖主观"感觉型"指标,难以产出有价值的数据。用 AI 来审查代码质量并不可靠:模型常常给出不一致的评分,或受命名等表面因素影响而改变偏好。尽管人工评估仍是保证可读性的金标准,但它缺乏现代训练流程和大规模基准所需的可扩展性。
为应对这一难题,研究者开始采用定量指标来区分 legacy codebases 与由 AI 生成的 slop 。两个较有前途的度量是 verbosity(通过启发式方法检测重复或不必要的冗长片段)和 erosion(衡量系统复杂度在多大程度上集中在少数过大、密集的函数中)。这些指标揭示了一个残酷事实:AI-generated code 在冗长度和侵蚀程度上约为 human-authored software 的两倍。
Agents 无法应对这种复杂性的现象在 SlopCodeBench 等基准中被凸显出来:这些基准通过在检查点间擦除模型上下文来模拟真实的迭代开发过程。在糟糕决策随时间累积的情形下,最先进的模型往往无法在多轮迭代中维持一个可用的 codebase 。对于那些在缺乏充分监督下大量引入 machine-generated code 的团队来说,这一失败是严重的警示。
归根结底,衡量代码质量仍离不开人类的直觉与审美。定量指标虽可帮助识别问题规模,但仅是迈向软件可维护性的一步。今后对 code churn 、 function cohesion 和 system coupledness 等指标的研究,将对完善这些工具至关重要。随着领域的发展,关注点必须从单纯生成代码,转向营造一个把结构完整性与功能性同等重视的开发环境。
While Large Language Models have become remarkably adept at generating formally correct code, their rise has ushered in a hidden crisis of code quality. Even when code passes automated tests, it often suffers from excessive abstraction, redundant duplication, and poor architectural decisions. This phenomenon, often termed slop, leads to an uncontrolled explosion of lines of code. As projects grow by millions of lines per month, human developers lose the ability to maintain oversight, and contrary to popular belief, AI agents are currently incapable of self-correcting this accumulation of technical debt.
The industry's current approach to evaluating this problem relies heavily on subjective, vibes-based metrics that fail to provide meaningful data. Using AI as a judge to grade code quality is largely ineffective, as models frequently produce inconsistent results or change their preferences based on superficial factors like naming conventions. Conversely, relying on human evaluation is the gold standard for maintaining readability, but it lacks the scalability required for modern training workflows and large-scale benchmarking.
To address this, researchers are turning to quantitative metrics to distinguish between legacy codebases and AI-generated slop. Two particularly promising measures are verbosity, which tracks duplicated or unnecessarily verbose segments using heuristics, and erosion, which calculates how much of a system's complexity is concentrated within a few oversized, dense functions. These metrics reveal a stark reality, as AI-generated code consistently proves to be roughly twice as verbose and eroded as human-authored software.
The inability of agents to manage this complexity is highlighted by benchmarks like SlopCodeBench, which simulate real-world, iterative development by erasing model context between checkpoints. Under these conditions, where bad decisions accumulate over time, state-of-the-art models often fail to maintain a functional codebase across multiple iterations. This failure serves as a critical warning for teams currently integrating vast amounts of machine-generated code into their systems without adequate oversight.
Ultimately, measuring code quality remains deeply tied to human intuition and taste. While quantitative metrics offer a way to identify the scale of the problem, they are only the beginning of a larger effort to ensure software remains manageable. Future explorations into code churn, function cohesion, and system coupledness will be essential to refining these tools. As the field evolves, the focus must shift from simply generating code to fostering an environment where structural integrity is treated with the same importance as technical functionality.
据报道,Iran-backed Houthi movement 已控制了 Perim Island,这座位于 Bab al-Mandab Strait 的小岛具有重要战略价值。该狭窄海峡连通 Red Sea 与 Gulf of Aden,是亚洲与欧洲之间的重要航运通道。夺岛属于其快速且大规模的军事攻势的一部分:此次行动已使 Houthi 控制了 Yemen 沿 Red Sea 的全部海岸线,迫使 Saudi-backed forces 撤退,并在一周内造成至少 46,000 人流离失所。 The Iran-backed Houthi movement has reportedly seized control of Perim Island, a strategically vital location situated within the Bab al-Mandab Strait. This narrow waterway serves as a critical gateway between the Red Sea and the Gulf of Aden, acting as an essential shipping corridor for trade between Asia and Europe. The capture of the island comes as part of a rapid, large-scale military offensive that has seen the Houthis take control of Yemen's entire Red Sea coast, forcing a retreat of Saudi-backed forces and displacing at least 46,000 people in just one week.
据报道,Iran-backed Houthi movement 已控制了 Perim Island,这座位于 Bab al-Mandab Strait 的小岛具有重要战略价值。该狭窄海峡连通 Red Sea 与 Gulf of Aden,是亚洲与欧洲之间的重要航运通道。夺岛属于其快速且大规模的军事攻势的一部分:此次行动已使 Houthi 控制了 Yemen 沿 Red Sea 的全部海岸线,迫使 Saudi-backed forces 撤退,并在一周内造成至少 46,000 人流离失所。
尽管 Houthi 领导层宣称行动取得成功并坚称国际海上通行安全,但他们同时明确表示对 Saudi 船只实施封锁。此举在 Riyadh 引发严重担忧:在 Strait of Hormuz 受阻后,Saudi Arabia 对经由 Red Sea 的石油出口通道愈发依赖。分析人士认为,Houthi 通过打击与 Saudi 有关的航运、同时向更广泛的国际社会释放安抚信号,可能是在有意向 Saudi Arabia 施压,同时尽量避免触发 Washington 的直接军事回应。
据报道,Saudi Crown Prince Mohammed bin Salman 已向 U.S. 请求直接军事干预,但 United States 拒绝直接出兵,只愿提供情报与目标指引支持,并将重心放在维护核心国家安全利益与保障航行自由上。 U.S. 官员称,他们正与地区盟友保持定期沟通,力图在应对这场升级危机与避免开辟新战线之间取得平衡。
这一冲突的战略重要性在 Djibouti 附近的重兵部署中凸显出来:包括 United States 、 China 和 France 在内的多国在当地设有基地,以监控并保护航道。随着 Yemen 沿海政府阵地突然崩溃、 Houthi 在关键通行区域站稳脚跟,观察人士警告局势高度不稳定。国际利益相关者最担忧的是全球能源市场可能进一步遭到冲击——若 Houthi 对这些海上咽喉的控制持续扩大,冲突将如何演变仍难以预测。
The Iran-backed Houthi movement has reportedly seized control of Perim Island, a strategically vital location situated within the Bab al-Mandab Strait. This narrow waterway serves as a critical gateway between the Red Sea and the Gulf of Aden, acting as an essential shipping corridor for trade between Asia and Europe. The capture of the island comes as part of a rapid, large-scale military offensive that has seen the Houthis take control of Yemen's entire Red Sea coast, forcing a retreat of Saudi-backed forces and displacing at least 46,000 people in just one week.
While the Houthi leadership claims that their military operations are a success and insists that international maritime navigation remains safe, they have explicitly stated that Saudi vessels are subject to a blockade. This move has caused significant alarm in Riyadh, as Saudi Arabia has become increasingly reliant on Red Sea routes for oil exports following the disruption of the Strait of Hormuz. Analysts suggest that by targeting Saudi-linked shipping while attempting to reassure the broader international community, the Houthis may be executing a calculated effort to pressure Saudi Arabia without triggering a direct military response from Washington.
In response to these developments, Saudi Crown Prince Mohammed bin Salman has reportedly sought direct U.S. military intervention. However, the United States has declined to engage directly, offering instead to provide intelligence and targeting support while maintaining a focus on protecting core national security interests and ensuring the freedom of navigation. U.S. officials maintain that they are in regular communication with regional allies, aiming to balance the management of this escalating crisis with an ongoing effort to avoid opening a new military front.
The strategic stakes of this conflict are underscored by the heavy military presence in nearby Djibouti, where several global powers, including the United States, China, and France, maintain bases to monitor and protect shipping lanes. With the sudden collapse of government-held positions along Yemen's coast and the Houthi entrenchment in key transit areas, observers warn that the situation remains highly volatile. The potential for further disruption to global energy markets remains a primary concern for international stakeholders, as it is difficult to predict how the conflict may evolve if Houthi influence over these vital maritime chokepoints continues to expand.
• Houthis 利用 deepfake 音频并在社交媒体上放大传播,制造混乱并诱使对方武装撤退,说明现代心理战能显著影响军事行动。
• 这类战术之所以有效,反映出部分地区部队缺乏严格的指挥链执行力;通过 Telegram 、 X 等具双重用途的应用进行的去中心化通信容易被用来散布虚假信息。
• 关于 AI 和 LLMs 在国防中的作用存在严重争议,担忧这些技术会降低恶意行为者实施复杂欺骗或设计武器的门槛。
• Houthis 阻断 Bab al-Mandab Strait 航运并袭击 Saudi 的关键基础设施,增强了其战略影响力,迫使全球重新评估石油供应的脆弱性。
• China 通过战略石油储备和电动汽车的快速普及,在一定程度上抵御了地区性能源冲击,而其他国家则更易受到供应链波动的影响。
• 关于近期 U.S. 干涉主义的动机与后果存在激烈分歧,批评者认为 Middle East 的政权更迭行动灾难性且无效,更多是受狭隘利益驱动而非追求长期稳定。
• 新闻报道风格的差异,尤其是标题中引号的使用,引发了有关偏见和"sane-washing"的指责;批评者认为西方媒体往往不成比例地将特定政府叙事包装成客观事实。
• Yemeni 冲突中对雇佣兵和非国家行为者的依赖造成了脆弱格局,士气低落、战术失误频发,破坏了地区权力结构的稳定。
• 关于主要 AI 实验室呼吁监管的动机是否出于切实的存在性担忧,还是为了维护对开源替代方案的竞争壁垒,各方仍存分歧。
• 具有核能力国家的扩散仍是全球地缘政治中的核心未决问题,关于以外交协议还是军事压力更能遏制意识形态政权,存在对立观点。
这场讨论反映了人们对快速技术进步与地缘政治稳定性退化交汇处的深切焦虑。参与者意见分歧:一方把这些事态归咎于情报机构的失败,另一方则视其为全球权力重新洗牌、更多是混乱化的表现。各方达成的明确共识是:对化石燃料的依赖以及对传统媒体和政府机构信任的丧失,创造了一个不稳定的环境,在这种环境中,即便是小规模的战术欺骗也可能产生不成比例的全球后果。最后,这场对话凸显了对军事与政治干预效果日益增长的愤世嫉俗,许多参与者对 U.S. 政策及媒体在报道这些危机时的透明度表示怀疑。
• The Houthis successfully leveraged a deepfake audio recording, amplified via social media, to sow confusion and induce a retreat among opposition forces, demonstrating the susceptibility of military operations to modern psychological warfare.
• The effectiveness of such tactics suggests a lack of disciplined chain-of-command enforcement in some regional forces, where decentralized communication via dual-use applications like Telegram and X can be exploited to spread disinformation.
• Significant debate exists regarding the role of AI and LLMs in national defense, with concerns that these technologies are lowering the barrier for bad actors to engage in sophisticated deception and weapon design.
• The Houthis' ability to disrupt shipping in the Bab al-Mandab Strait and strike critical Saudi infrastructure has elevated their strategic impact, forcing a re-evaluation of global oil supply vulnerabilities.
• China's strategic oil stockpiling and rapid adoption of electric vehicles have partially insulated its economy from regional energy shocks, whereas other nations remain more exposed to supply chain volatility.
• There is intense disagreement over the motivation and outcomes of recent U.S. interventionism, with critics arguing that regime-change operations in the Middle East have been catastrophic, ineffective, and driven by narrow interests rather than long-term stability.
• Discrepancies in journalistic reporting styles, specifically the use of quotation marks in headlines, have sparked accusations of bias and "sane-washing," with critics arguing that Western media outlets disproportionately platform specific government narratives as objective fact.
• The reliance on mercenaries and non-state actors in the Yemeni conflict has created a fragile landscape where morale is low and tactical failures are common, undermining the stability of regional power structures.
• Disagreement persists regarding the extent to which major AI labs' calls for regulation are motivated by genuine existential concerns versus a desire to protect their competitive moats against open-source alternatives.
• The proliferation of nuclear-capable rogue states remains a central, unresolved tension in global geopolitics, with opposing views on whether diplomatic deals or military pressure are more effective at containing ideological regimes.
The discussion reflects a deep anxiety regarding the intersection of rapid technological advancement and the degradation of geopolitical stability. Participants are divided between views that see these developments as a failure of institutional intelligence and those who interpret them as part of a broader, chaotic rebalancing of global power. There is a clear consensus that the reliance on fossil fuels, combined with the loss of trust in traditional media and government institutions, has created a volatile environment where small-scale tactical deceptions can have disproportionate, global consequences. Ultimately, the conversation highlights a growing cynicism toward the efficacy of military and political interventions, with many contributors expressing skepticism toward both U.S. policy and the transparency of the media reporting on these crises.
"The Waymo effect"描绘了我们与技术以及彼此互动方式的一种悄然变化。通过消除与他人打交道的摩擦,技术带来一种看似纯粹的收益感。例如,我们欣赏无人驾驶汽车的便利,因为它让我们绕过人类互动中不可避免的寒暄与社交摩擦。但这种便利也有隐性代价:我们失去了那些非自选且往往不可预测的与陌生人相遇的机会,而正是这些相遇能拓宽视野,把我们与自身信息泡沫之外的世界连接起来。 The Waymo effect describes a quiet shift in how we interact with technology and each other. By removing the friction of dealing with other people, technology offers an experience that feels like pure gain. We appreciate the convenience of the driverless car, for example, because it allows us to bypass the small talk and social friction inherent in human interaction. Yet, this convenience comes at a hidden cost. We lose the unchosen, often unpredictable encounters with strangers that broaden our perspectives and connect us to a world outside our own bubbles.
"The Waymo effect"描绘了我们与技术以及彼此互动方式的一种悄然变化。通过消除与他人打交道的摩擦,技术带来一种看似纯粹的收益感。例如,我们欣赏无人驾驶汽车的便利,因为它让我们绕过人类互动中不可避免的寒暄与社交摩擦。但这种便利也有隐性代价:我们失去了那些非自选且往往不可预测的与陌生人相遇的机会,而正是这些相遇能拓宽视野,把我们与自身信息泡沫之外的世界连接起来。
这种动力正在逐步重塑学术与科学研究的格局。正如无人驾驶汽车一样,大型语言模型为人类合作者提供了一种无摩擦、随时可得的替代方案。同行可能带来令人不便的议程、分歧或拒绝配合你的前提,但人工智能助手则完全顺从、随时可用、不需妥协。然而,人类合作者的"麻烦"恰恰往往蕴含真正价值:正是在这种摩擦中,想法被挑战、假设被检验、意外的机遇才会出现。当研究者用聊天机器人的无缝产出取代人类之间那种混乱且不可预测的协作过程时,就存在一种协作瓦解的风险,这会威胁到研究共同体的社会结构基础。
现代学术界的激励机制进一步加剧了这一倾向。机构与资助体系在很大程度上更看重速度与可衡量的产出,使得在许多情况下选择机器而非人类成为最理性的决策。协作本身耗时且昂贵,涉及出差、建立信任以及管理参与者的自尊心。当预算吃紧、职业晋升依赖于迅速发表时,缓慢而深入的人类协作很容易被边缘化。结果是,个人生产力可能上升,但思想的集体多样性却缩水——人人依赖那些能产出流畅、结构化但往往趋同结论的工具。
此外,写作不仅是产出文本的行为,更是思考的重要组成部分。写作像一种强制机制,把模糊的直觉锻造成精确的论点,揭示出可能被忽略的逻辑漏洞。把写作过程外包给 AI,研究者可能无意中跳过那些困难而必要的深度思考,最终用表面光鲜、无摩擦的模仿替代了扎实的理解。危险在于,学术文化表面上看起来健康且高效,而支撑它的关键而无形的智力劳动却在悄然流失。
归根结底,这并不是要拒绝新技术——它们确实带来创新与效率的真实潜力。相反,这是提醒我们正视一个问题:当机器驱动的工作流程变得廉价、而以人为驱动的协作变得非常昂贵时,系统本身存在风险。若要维护研究的质量与深度,就必须有意识地把摩擦重新引入体系:将非结构化的人际互动视为关键的基础设施,奖励研究者的协作贡献,并确保人类牢牢掌握主导地位。在论文撰写能力变得普及之后,真正稀缺且有价值的,将是共同思考的能力。
The Waymo effect describes a quiet shift in how we interact with technology and each other. By removing the friction of dealing with other people, technology offers an experience that feels like pure gain. We appreciate the convenience of the driverless car, for example, because it allows us to bypass the small talk and social friction inherent in human interaction. Yet, this convenience comes at a hidden cost. We lose the unchosen, often unpredictable encounters with strangers that broaden our perspectives and connect us to a world outside our own bubbles.
This dynamic is increasingly shaping the landscape of academic and scientific research. Much like the driverless car, large language models provide a frictionless, always-available substitute for the human collaborator. While a colleague might arrive with inconvenient agendas, disagreements, or the refusal to align with your specific assumptions, an AI assistant is perfectly compliant, available at any hour, and requires no compromise. However, the very inconvenience of a human collaborator is often where the real value lies. It is through this friction that ideas are challenged, assumptions are tested, and serendipity occurs. When researchers trade the messy, unpredictable process of human collaboration for the seamless output of a chatbot, they risk a process of decollaboration that threatens the underlying social fabric of the research community.
The incentive structures within modern academia further exacerbate this trend. Institutions and funding systems largely prioritize speed and measurable output, creating a environment where choosing a machine over a human is often the most rational decision. Collaboration is inherently time-consuming and expensive, involving travel, trust-building, and the management of egos. When budgets are tight and career advancement depends on rapid publication, the slow, human work of deep collaboration is easily sidelined. This creates a trap where individual productivity may rise, but the collective diversity of ideas narrows, as everyone relies on tools that provide fluent, structured, yet often convergent results.
Furthermore, we must recognize that writing is not just the production of an artifact, but an essential component of thinking. It serves as a forcing function that turns vague intuitions into precise arguments, revealing holes in logic that might otherwise go unnoticed. By outsourcing the writing process to AI, researchers may inadvertently skip the difficult, effortful thinking required for genuine insight. This transition risks replacing durable understanding with a polished, frictionless imitation. The danger is that research culture could appear healthy and efficient on the surface while the vital, invisible intellectual labor that sustains it slowly erodes.
Ultimately, this is not an argument for rejecting new technologies, which offer genuine potential for innovation and efficiency. Instead, it is a call to recognize the risks of a system that makes machine-driven workflows cheap and human-driven collaboration prohibitively expensive. If we wish to preserve the quality and depth of research, we must deliberately engineer friction back into the system. This means valuing unstructured, human interaction as a form of critical infrastructure, rewarding researchers for their collaborative contributions, and ensuring that humans remain firmly in the driver's seat. As the ability to produce papers becomes common, the truly scarce and valuable resource will be the ability to think together.
• 非专业人员倾向于把 LLMs 视为"真理来源"来挑战专业判断,这造成严重摩擦:人们往往更信任 AI 的输出,而不是经验丰富的合作者的细致判断。
• 对 LLMs 的依赖催生了"meat proxy"现象:团队成员给出的复杂但误导性的论证,反而更像是与 AI 交互的拼凑产物,而非独立的批判性思维。
• 各领域专家越来越感到沮丧,他们不得不反复向那些把 LLM 的自信当作领域专业替代品的客户或同事,证明基本决策的合理性。
• 出现了一种明显的行为模式:个别人借助 LLMs 绕过深入学习的必要性,以高速且过于自信的方式实施方案,往往缺乏对权衡或基本约束的理解。
• "Physics Graduate" 这一比喻描绘了一类典型人物:进入新领域时表现出傲慢和缺乏谦逊的高智商个体,常常忽视既有专业知识,用拙劣的方法重造系统。
• 怀疑精神和对 AI 输出进行验证的能力,是有效从业者的基本素质;这与一类用户形成鲜明对比——他们缺乏批判能力,把 LLM 的结论奉为圭臬。
• 追求"无摩擦"生活(即通过技术消除对不可预测性和他人社交需求的应对)可能正在加剧社会孤立,并削弱协作研究与问题解决的能力。
• 通过 AI 实现的信息"民主化"正受到质疑:这种商品化模式用综合处理取代了真正理解一门学科所需的认知劳动。
• 辨别文本是否由 AI 生成成了常见争论点,批评者认为一些"Claude-isms"和特定修辞结构暴露了缺乏真正人类创作的迹象。
• 根本挑战在于作为 AI 工具的"conductor"保持自主权;长期风险在于,引导这些系统所需的深度、领域特定知识可能会逐渐萎缩。
此次讨论反映出两个紧张的方面:AI 作为生产力工具的效用,与它作为知识孤立和傲慢催化剂的潜在能力。共识是:尽管 AI 可加速工作流程,但当用户缺乏承认自身知识局限和对机器权威保持谦逊时,AI 往往会侵蚀协作过程。这里识别出的行为模式表明,我们正朝向一种优先追求"即时答案"的便捷文化,而非愿意承受现实世界对话和专家审查带来的摩擦。参与者认为,除非从业者继续致力于基础性探究并坚持高标准的验证,否则对 AI 的依赖可能导致质量下降、组织记忆流失,以及人类在专业环境中互动方式的深刻改变。
• The rising tendency for non-experts to use LLMs as a "source of truth" to challenge professional expertise creates significant friction, as individuals often prioritize AI-generated output over the nuanced judgment of experienced collaborators.
• Reliance on LLMs has created a "meat proxy" phenomenon, where team members deliver sophisticated but misguided arguments that appear to be synthesized fragments of AI interactions rather than independent critical thinking.
• Experts across various fields report increasing frustration with being forced to repeatedly justify fundamental decisions to clients or colleagues who view LLM confidence as a substitute for domain-specific mastery.
• A clear pattern of behavior has emerged where individuals use LLMs to bypass the need for deep learning, trading genuine expertise for high-speed, overconfident implementation that often lacks a grasp of tradeoffs or underlying constraints.
• The "Physics Graduate" trope identifies a specific archetype of high-intelligence individuals who exhibit arrogance and a lack of humility when entering new domains, often ignoring established expertise in favor of re-inventing systems poorly.
• Skepticism and the ability to verify AI outputs are essential traits of effective practitioners, contrasting with a growing tier of users who accept LLM results as gospel because they lack the ability to critique the output.
• The quest for a "frictionless" life—where technology removes the need to deal with the unpredictability and social requirements of other humans—may be leading toward increased social isolation and a degradation of collaborative research and problem-solving.
• The perceived "democratization" of information through AI is contested as a form of commodification, where synthesis replaces the cognitive work required to actually understand a subject.
• Discerning whether a text is AI-generated has become a common point of contention in discourse, with critics arguing that certain "Claude-isms" and rhetorical structures signal a lack of genuine human authorial effort.
• Ultimately, the challenge lies in maintaining agency as a "conductor" of AI tools, as the long-term risk involves the atrophy of the deep, domain-specific knowledge required to guide these systems effectively.
The discussion reflects a widespread tension between the utility of AI as a productivity tool and its capacity to act as a catalyst for intellectual isolation and arrogance. A consensus emerges that while AI can accelerate workflows, it frequently erodes the collaborative process when users lack the humility to recognize the limits of their own knowledge or the authority of the machine. Patterns of behavior identified here suggest that we are moving toward a culture where the convenience of "instant answers" is prioritized over the friction of real-world dialogue and expert vetting. The participants argue that unless practitioners remain committed to fundamental inquiry and maintain high standards for verification, the reliance on AI will likely lead to a decline in quality, the loss of shared organizational memory, and a problematic shift in how humans interact within professional environments.
RTK,即 Rust Token Killer,是一款通过过滤和压缩终端输出以降低 AI 编码成本的工具,近来颇为流行。尽管有许多病毒式传播的帖子声称启用 RTK 可将 Token 使用量减少高达 90%,但这类说法常把终端数据量的减少等同于实际费用的节省。对 Claude Code 和 DeepSeek V4 Pro 等模型进行了超过 1,700 次的严格基准测试后,结果表明 RTK 并不可靠地降低总体费用,甚至在某些情况下会提高成本。 RTK, or Rust Token Killer, has gained significant popularity as a tool designed to reduce AI coding costs by filtering and compressing terminal output. While many viral posts and claims suggest that implementing RTK can cut token usage by up to 90%, these projections often conflate reduced terminal volume with actual cost savings. After conducting a rigorous benchmark test across more than 1,700 attempts using models like Claude Code and DeepSeek V4 Pro, the findings indicate that RTK does not reliably lower total expenses and can, in some cases, increase them.
RTK,即 Rust Token Killer,是一款通过过滤和压缩终端输出以降低 AI 编码成本的工具,近来颇为流行。尽管有许多病毒式传播的帖子声称启用 RTK 可将 Token 使用量减少高达 90%,但这类说法常把终端数据量的减少等同于实际费用的节省。对 Claude Code 和 DeepSeek V4 Pro 等模型进行了超过 1,700 次的严格基准测试后,结果表明 RTK 并不可靠地降低总体费用,甚至在某些情况下会提高成本。
测试采用 Terminal-Bench 2.1 在终端交互频繁的场景中评估工具表现。结果显示,在使用 Fable 5.0 的 Claude Code 的某个特定场景下,成本仅略降约 5%;而在 DeepSeek 上,成本反而上升了 5% 。更重要的是,全部任务的平均值显示,Fable 并未呈现明确的成本优势,而 DeepSeek 的开销平均上升了 17% 。可见,RTK 报告的"收益"是基于丢弃的字节数而非计费的 Token 数,这一指标容易误导对实际财务节省的判断。
测试中还发现 RTK 容易引入技术性错误。一次插件故障导致代理进入死循环,出现了 339 次连续错误,使得总成本大约是基线的九倍。此外,现代 AI 代理通常已经具备管理输出的手段(例如使用 head 或 tail 命令),因此来自终端输出的账单占比本就不高。 RTK 的压缩效果常被抵消,因为压缩终端输出后,代理往往需要额外的交互轮次才能完成任务,从而抹去了潜在的节省。
综上,研究认为 RTK 更像是一种针对特定场景的小众优化工具,而非普适的 AI 编码成本解决方案。代理式编码模型高度依赖缓存读取,处理原始终端输出的成本本就比开发者预想的低得多。鉴于额外的代理轮次代价较高,使用压缩工具通常难以带来净收益。随着前沿模型在终端交互处理方面变得越来越高效,RTK 在普遍意义上的成本下降效果仍未得到证实。
RTK, or Rust Token Killer, has gained significant popularity as a tool designed to reduce AI coding costs by filtering and compressing terminal output. While many viral posts and claims suggest that implementing RTK can cut token usage by up to 90%, these projections often conflate reduced terminal volume with actual cost savings. After conducting a rigorous benchmark test across more than 1,700 attempts using models like Claude Code and DeepSeek V4 Pro, the findings indicate that RTK does not reliably lower total expenses and can, in some cases, increase them.
The testing process utilized Terminal-Bench 2.1 to evaluate how the tool performs in environments with heavy terminal interaction. Results showed that for Claude Code using Fable 5.0, costs decreased by a marginal 5% in one specific scenario, while for DeepSeek, costs actually increased by 5%. More importantly, when averaged across all tasks, Fable saw no clear cost benefit, and DeepSeek's expenses rose by 17% on average. These results demonstrate that the tool's reported gain—which measures discarded bytes rather than billed tokens—is a misleading metric for estimating financial savings.
A significant issue identified during testing is the propensity for RTK to introduce technical errors. In one instance, a plugin bug caused an agent to enter an infinite loop, resulting in 339 consecutive errors and a total cost roughly nine times higher than the baseline. Furthermore, because modern AI agents already possess built-in techniques for managing output, such as using head or tail commands, the actual proportion of an agent's bill derived from terminal output is relatively small. The compression offered by RTK is frequently offset by the fact that agents often require additional turns to complete a task when the terminal output is condensed, effectively erasing any potential savings.
Ultimately, the study concludes that RTK is a niche optimization tool rather than a comprehensive solution for reducing AI coding costs. Because agentic coding models rely heavily on cache reads, the cost of processing raw terminal output is already significantly lower than developers might assume. Since extra agent turns are costly, the trade-off of using a compression tool rarely results in a net gain. Given that frontier models are becoming increasingly efficient at handling terminal interactions on their own, the practical utility of RTK for general cost reduction remains unproven.
• 局部语义搜索,例如用 embedding 模型对代码库建立索引,让 LLMs 根据相似度而非穷举式字符串搜索来识别相关代码段,这比蛮力的 grep 命令更有潜力作为有效替代方案。
• 许多面向 LLMs 的病毒式"生产力技巧",例如 RTK 或各种命令输出压缩器,越来越被视为无效甚至有害,独立评估通常不能证明它们在性能或成本效益上有统计学显著的提升。
• 这些工具作者给出的基准结果常常不可复现,且经常依赖误导性指标,比如基于理论输入字节数而非模型实际行为或端到端任务表现来计算"节省"。
• "caveman" 提示技术被一些人认为是一种既实用又带点幽默感的强制简洁手段,可能通过减少会削弱模型推理的指令臃肿来提高准确性。
• 强制工具对 CLI 输出进行压缩或过滤常会令 LLMs 困惑,因为这些模型是在标准 shell 行为下预训练的,面对非标准、被截断或被摘要化的输出时,可能误以为工具已损坏或返回了不完整的数据。
• 使用 IDE-native 工具(例如 JetBrains MCP servers)通常比通用命令行包装器更稳健,因为它们公开的是高级操作,而不是让 LLM 解析原本不可预测的终端输出。
• 节省 token 最有效的办法往往很简单:在标准构建工具中利用现有的选项(如 --quiet 或 silent 模式),而不是引入可能造成错误并增加延迟的复杂中间件。
• 将自动汇总大型工具输出的任务卸载给更便宜、更小的模型(例如 flash 或 lite 变体),是在不牺牲准确性或破坏 agent 逻辑的前提下保持上下文的更可靠做法。
• 人们对这些工具那种"凭感觉"式的营销持强烈怀疑,不少开发者认为配置和调试这些技巧所耗的时间通常超过了任何微小且未经证实的 token 成本节约。
• 归根结底,基于 agent 的开发最可靠的策略是让模型与预期的、标准化的输出交互,并仅用简单、透明且明确限制输出量的脚本进行干预,而不是通过不透明的抽象层去操纵输出。
有经验的开发者之间普遍达成的共识是:许多针对 LLM coding agents 的"生产力提升"技巧在很大程度上是虚幻产品,带来的问题往往多于解决的问题。通过干扰模型被训练用来解读的标准输出,这些工具经常削弱模型的推理能力,迫使 agents 进入重试循环,从而抵消任何潜在的 token 节省。追求效率是合理的,但最有效的策略通常是与模型原生能力相配合的做法,例如使用现有 CLI 的冗长度选项或把摘要任务卸载给更便宜的模型,而不是依赖所谓的"魔法"包装器。开发者正越来越批判性地要求独立验证,并倾向于将稳定的原生集成置于那些未经证实且常带误导性的优化工具之上。
• Local semantic search, such as indexing codebases with embedding models, offers a potentially more efficient alternative to brute-force grep commands by allowing LLMs to identify relevant code sections through proximity rather than exhaustive string searching.
• Many viral "productivity hacks" for LLMs, such as RTK or various command-output compressors, are increasingly viewed as ineffective or even detrimental, often failing to demonstrate statistically significant improvements in performance or cost-efficiency in independent evaluations.
• Benchmark results provided by authors of these tools are frequently unreproducible and often rely on misleading metrics, such as calculating "savings" based on theoretical input bytes rather than actual model behavior or end-to-end task performance.
• The "caveman" prompt technique is viewed by some as a practical, albeit humorous, method to enforce conciseness, which potentially improves accuracy by reducing the instruction bloat that can degrade model reasoning.
• Forcing tools to "compress" or filter CLI output frequently confuses LLMs, as they are pre-trained on standard shell behavior and may incorrectly assume a tool is broken or return incomplete data when faced with non-standard, truncated, or summarized output.
• Using IDE-native tools—like JetBrains MCP servers—is considered a more robust approach than generic command-line wrappers, as they expose high-level actions rather than relying on the LLM to parse raw, unpredictable terminal output.
• The most effective way to save tokens is often the simplest: utilizing existing flags like `--quiet` or `silent` modes in standard build tools, rather than introducing complex middleware that can induce errors and increase latency.
• Automating the summarization of large tool outputs by offloading the task to a cheaper, smaller model (such as a flash or lite variant) is a more reliable way to maintain context without sacrificing accuracy or breaking the agent's logic.
• There is a strong skepticism regarding the "vibe-coded" marketing of these tools, with many developers concluding that the time spent configuring and debugging these hacks often outweighs any marginal, unverified token cost reduction.
• Ultimately, the most reliable strategy for agent-based development is to allow the model to interact with expected, standard outputs and intervene only with simple, transparent scripts that explicitly limit output volume rather than manipulating it through opaque abstraction layers.
The consensus among experienced developers suggests that many current productivity "hacks" for LLM coding agents are largely vaporware that create more problems than they solve. By interfering with the standard outputs that models were trained to interpret, these tools frequently degrade reasoning and force agents into retry loops that negate any potential token savings. While the search for efficiency is legitimate, the most effective strategies appear to be those that work with the model's native capabilities, such as using existing CLI verbosity flags or offloading summarization to cheaper models, rather than employing proprietary "magic" wrappers. Developers are increasingly moving toward a more critical stance, demanding independent verification and prioritizing stable, native integrations over unproven and often misleading optimization tools.
Claude 仅供年满 18 岁的用户使用,用户在创建账户时必须确认自己符合该年龄要求。平台使用自动化安全系统检测可能未满 18 岁的迹象;若系统怀疑有未成年人使用某账户并将其标记,访问权限会被限制,用户需先完成年龄验证后方可继续使用服务。 Claude is strictly intended for users who are at least 18 years old, and individuals must confirm they meet this age requirement during the initial account setup process. The platform employs automated safety systems designed to detect indicators that a user might be under the age of 18. If the system flags an account due to suspected minor activity, access will be restricted, and the user will be prompted to verify their age before they can continue using the service.
Claude 仅供年满 18 岁的用户使用,用户在创建账户时必须确认自己符合该年龄要求。平台使用自动化安全系统检测可能未满 18 岁的迹象;若系统怀疑有未成年人使用某账户并将其标记,访问权限会被限制,用户需先完成年龄验证后方可继续使用服务。
为便于验证,平台与独立第三方年龄验证服务商 Yoti 合作。账户被标记时,用户会收到一封包含验证链接的通知邮件;成功完成验证后,账户将恢复使用。
用户可通过 Yoti 选择多种方式确认年龄:可以选择面部年龄评估,通过自拍估算年龄,无需提交证件;也可以选择上传政府签发证件的照片进行 ID 验证,例如 Passport 或 Driver's License;已使用 Yoti Digital ID App 的用户则可直接从 App 分享已验证的 "over 18" 属性。
在数据隐私方面,Yoti 是一家经独立审计且符合 SOC2 的提供商,保障安全。验证流程旨在保护用户隐私——平台不会查看或存储任何身份证件或自拍,Yoti 在年龄核验完成后会立即删除所有个人数据、文档和图像。平台仅接收通过或未通过的结果,确保验证过程中不保留任何敏感个人信息。
Claude is strictly intended for users who are at least 18 years old, and individuals must confirm they meet this age requirement during the initial account setup process. The platform employs automated safety systems designed to detect indicators that a user might be under the age of 18. If the system flags an account due to suspected minor activity, access will be restricted, and the user will be prompted to verify their age before they can continue using the service.
To facilitate this verification, the platform partners with Yoti, an independent, third-party age verification provider. When an account is flagged, the user receives a notification email containing a link to begin the verification process. Successfully completing this process allows the user to have their account reinstated.
Users have multiple options for confirming their age through Yoti. They may choose facial age estimation, which uses technology to estimate their age from a selfie without requiring an ID document. Alternatively, they can opt for ID verification by uploading a photo of a government-issued document, such as a passport or driver's license. Finally, those who already use the Yoti digital ID app can share a verified "over 18" attribute directly from the app.
Regarding data privacy, Yoti acts as an independently audited, SOC2-compliant provider to ensure security. The verification process is designed to protect user privacy, as the platform does not see or store any ID documents or selfies. Yoti deletes all personal data, documents, and images immediately after the age check is completed. Consequently, the platform only receives a pass or fail result, ensuring that no sensitive personal information from the verification process is stored by the company.
• AI 公司实施年龄验证的决定普遍被视为一种策略性举措,用以减轻与未成年人相关的法律与监管风险,而非真正出于保护儿童或提升安全性的考虑。
• 很多人认为年龄验证对隐私构成重大威胁,担心第三方供应商收集并存储身份证件会形成一个庞大而脆弱的"蜜罐",从而增加数据泄露的风险。
• 一个反复出现的观点是,采用限制性访问的方式对儿童和青少年进行非人化对待——批评者认为这是一种矫枉过正,剥夺了年轻人获取学习、创造和成长所需工具的机会。
• 一些参与者认为应由父母而非公司或政府来监督和管理子女的互联网访问,他们更倾向于依靠家长控制而非强制性的身份验证来解决问题。
• 对年龄验证有效性的怀疑仍然存在,许多人指出这不仅为合法用户设置了不必要的障碍,也难以阻止有动机的未成年人绕过这些控制措施。
• 有人对"为了孩子着想"的说辞表示不满,认为这是公司为逃避更深层次系统性改革或为收集更详尽的用户数据寻找的方便借口。
• 舆论普遍认为,高质量、自托管、开放权重模型的兴起,为希望规避侵入性身份验证的用户提供了可行且更能保护隐私的替代方案。
• 有人建议采用操作系统层面的标记或基于顶级域名(TLD)的过滤等技术替代方案,以较不侵入的方式处理年龄限制,但也有人担心这些功能可能被滥用,用于更广泛的监控或国家控制。
• 讨论反映出对大型科技公司的根深蒂固不信任,许多用户认为利润驱动和法律免责始终凌驾于用户自主权与隐私之上。
• 关于 AI 对青少年的长期影响是否本质上有害存在显著争议:有人把 AI 看作现代学习不可或缺的"数字导师",而另一些人则担心其可能导致成瘾、发育迟缓或心理伤害。
这场讨论的核心在于:人们对 AI 作为年轻人创造和教育工具的感知效用,与科技公司日益承受的法律和监管压力之间存在根本张力。大多数参与者对 AI 提供商的动机持怀疑态度,认为年龄验证是一种"合规表演",既把法律责任推给父母,又为进一步收集用户数据制造借口。尽管部分人承认未成年人确有潜在风险,但普遍情绪是对数字自由被侵蚀的担忧,以及对互联网正朝着以身份验证为前提的方向发展的沮丧。许多用户认为,转向本地开源模型是对这些侵犯隐私行为的合乎逻辑反应,标志着在面对集中化企业控制时,人们正朝自给自足的方向转变。
• The decision by AI companies to implement age verification is widely perceived as a tactical move to mitigate legal and regulatory liability regarding minors, rather than a genuine effort to protect children or improve safety.
• Many see age verification requirements as a significant privacy threat, fearing that the collection and storage of identity documents by third-party vendors creates a massive, vulnerable honeypot for data breaches.
• A recurring theme is the perceived dehumanization of children and teenagers through restrictive access, which critics argue is an overcorrection that deprives young people of valuable tools for learning, creativity, and development.
• Several participants argue that parents, not corporations or governments, should maintain the authority to oversee and manage their children's internet access, favoring parental controls over mandatory identity checkpoints.
• Skepticism persists regarding the efficacy of age verification, with many noting that it creates unnecessary friction for legitimate users while failing to stop minors who are motivated enough to circumvent the controls.
• Some express frustration at the "think of the children" rhetoric, characterizing it as a convenient excuse for companies to avoid deeper systemic reforms or to justify the collection of more granular user data.
• There is a strong consensus that the rise of high-quality, self-hosted, open-weight models provides a viable, privacy-preserving alternative for users who wish to avoid intrusive ID verification requirements.
• Technical alternatives, such as OS-level flags or TLD-based filtering, were proposed as less invasive ways to handle age-gating, though concerns were raised about the risk of such features being abused for broader surveillance or state control.
• The discussion reflects a deep-seated distrust of big tech companies, with many users asserting that profit incentives and liability protection always take precedence over user agency and privacy.
• Significant debate exists over whether the long-term impact of AI on youth is inherently negative, with some framing it as a "digital tutor" essential for modern learning, while others worry about the potential for addiction, stunting, or psychological harm.
The conversation centers on a fundamental tension between the perceived utility of AI as a creative and educational tool for young people and the increasing legal and regulatory pressure on companies to restrict access. Most participants are cynical about the stated motivations of AI providers, viewing age verification as "compliance theater" designed to shift legal liability onto parents while simultaneously gathering more data on the user base. While some acknowledge the reality of potential risks to minors, the prevailing sentiment is one of frustration toward the erosion of digital freedom and the move toward an internet where access is contingent upon identity verification. Many users suggest that the shift toward local, open-source models is the logical reaction to these privacy-invasive practices, signaling a broader movement toward self-sufficiency in the face of centralized corporate control.
该文档只是一个自动安全验证页面,用来防止恶意机器人流量,表明网站正在通过标准的 Cloudflare 安全检查来核实用户身份。 The provided document serves only as an automated security verification page used to protect the website from malicious bot activity. It indicates that the site is currently confirming the user's legitimacy through a standard Cloudflare security check.
该文档只是一个自动安全验证页面,用来防止恶意机器人流量,表明网站正在通过标准的 Cloudflare 安全检查来核实用户身份。
由于输入仅包含该技术验证通知、并不包含实际文章内容,因此无法就具体内容、论点或主题作出摘要。该页面只是通往目标站点的通道,并不包含实质性信息。
The provided document serves only as an automated security verification page used to protect the website from malicious bot activity. It indicates that the site is currently confirming the user's legitimacy through a standard Cloudflare security check.
Because the input consists exclusively of this technical verification notification and lacks the actual article text, it is not possible to generate a summary of specific content, arguments, or thematic details. The page serves as a gateway to the intended site rather than containing substantive information.
• Cherenkov radiation 发生在带电粒子在某种介质中运动速度超过该介质中的光的相速度时,产生类似音爆的效应。
• "光速"这一术语经常被误解,常与基本常数 c(因果速度 / 真空中的光速)混淆,而后者仍是不可逾越的宇宙极限。
• Imaging Atmospheric Cherenkov Telescopes (IACTs) 以及像 Pierre Auger Observatory 和 IceCube 这样的大型水或冰质探测器,利用 Cherenkov radiation 特有的蓝色闪光来探测并重建高能宇宙射线和中微子。
• 光在介质中的表观变慢并不是因为单个光子真的减速或发生非相干弹射,而是一种复杂的集体电磁效应,涉及介质原子产生的相位移动与再辐射波的干涉。
• 把 Cherenkov radiation 比作"音爆"有助于形象化,但并不完整。关键区别在于,粒子速度超过局部相速度会造成相长干涉,把发射的光汇聚成一束相干且有方向性的锥形辐射。
• 20 世纪 60 年代为提高 Cherenkov 探测器的光收集效率而进行的研究催生了非成像光学(non-imaging optics)的发展,该技术如今已广泛应用于太阳能和照明设计。
• 一些宇航员报告说,由于宇宙射线与眼内玻璃体相互作用,他们曾在眼中直接观察到 Cherenkov 闪光,这为该现象提供了直观的亲身体验。
• 在某些历史性核事故(例如 Goiânia 事故)中观察到的蓝色辉光仍在研究中,尽管通常被归因为大气电离或荧光,而非纯粹的 Cherenkov radiation 。
• 关于"超光速旅行"标题是否误导的讨论凸显了科学传播中的常见张力:技术准确性常与追求吸引眼球的简化解释发生冲突。
• Cherenkov radiation 作为抽象物理学与可观测的、近乎"魔法"般现象之间的桥梁,提供了通过有形的高能事件来见证物理定律极限的独特窗口。
本次对话澄清了围绕"光速"一词的长期混淆,通过区分常数 c 与在折射介质中低得多的相速度,达成了广泛共识:Cherenkov radiation 并不违背相对论,而是一种电磁冲击波,是现代天体物理学中观测那些原本不可见的高能粒子的关键工具。尽管对该现象的解释可以从简单类比延展到复杂的波传播理论,但这一话题共同强调了一个认识:光、物质与速度之间的相互作用为我们洞察基本物理过程提供了重要窗口。
• Cherenkov radiation occurs when charged particles travel through a medium at a speed greater than the phase velocity of light in that same medium, creating an effect analogous to a sonic boom.
• The term "speed of light" is frequently misunderstood because it is often conflated with the fundamental constant c (the speed of causality/speed of light in a vacuum), which remains an unbreakable universal limit.
• Imaging Atmospheric Cherenkov Telescopes (IACTs) and large-scale water or ice detectors like the Pierre Auger Observatory and IceCube utilize the characteristic blue flash of Cherenkov radiation to detect and reconstruct high-energy cosmic rays and neutrinos.
• The apparent slowing of light in a medium is not due to individual photons physically slowing down or bouncing incoherently; rather, it is a complex collective electromagnetic effect involving phase shifts and re-emitted waves from the medium's atoms.
• The "sonic boom" analogy for Cherenkov radiation is helpful for visualization but scientifically incomplete; the key distinction is that moving faster than the local phase velocity allows for constructive interference, focusing the emitted light into a coherent, directional cone.
• Research into efficient light collection for Cherenkov detectors in the 1960s led to the development of non-imaging optics, which now has broad applications in solar energy and illumination design.
• Astronauts have reported observing Cherenkov flashes directly in their eyes due to cosmic ray interactions with the vitreous humor, providing a visceral, personal experience of the phenomenon.
• The blue glow observed in some historical nuclear accidents, such as the one in Goiânia, remains a subject of investigation, though it is often attributed to atmospheric ionization or fluorescence rather than pure Cherenkov radiation.
• Discussions regarding the "misleading" nature of headlines about faster-than-light travel highlight a common tension in science communication, where technical accuracy often conflicts with the desire for engaging, simplified explanations.
• The phenomenon of Cherenkov radiation serves as a bridge between abstract physics and observable, "magic-like" reality, providing a unique opportunity to witness the limits of physical laws through tangible, high-energy events.
The conversation clarifies the persistent confusion surrounding the term "speed of light" by distinguishing between the constant c and the slower phase velocity of light within a refractive medium. There is a strong consensus that Cherenkov radiation is not a violation of relativity, but rather an electromagnetic shockwave that acts as a vital tool in modern astrophysics for observing otherwise invisible high-energy particles. While the technical explanations range from simple analogies to complex wave-propagation physics, the thread underscores a shared appreciation for how these interactions between light, matter, and velocity offer a window into fundamental physical processes.
当前 AI 工程越来越呈现内卷化倾向:系统要求投入更多努力、产出更高竞争性成果,却并未真正提升价值或生产力。虽然像 GPT 6 Astra 这样的新模型在处理计算机操作、图像理解和复杂任务方面能力突出,但它们给实际的软件工程带来了显著挑战。模型似乎缺乏明确的约束或惩罚机制来抑制低质量代码的生成,因而对长期任务完成的过度关注反而催生出大量不可用或不可维护的产出。 The current state of AI engineering is increasingly characterized by what can be described as Neijuan, or involution, where systems demand greater effort and competitive output without any real increase in value or productivity. While new models like GPT 6 Astra are undeniably impressive in their capacity to handle computer use, image understanding, and complex tasks, they present significant challenges for practical software engineering. These models seem to lack clear constraints or penalties for producing low-quality code, leading to an environment where the focus on long-horizon task completion results in the creation of vast amounts of unusable or unmaintainable material.
当前 AI 工程越来越呈现内卷化倾向:系统要求投入更多努力、产出更高竞争性成果,却并未真正提升价值或生产力。虽然像 GPT 6 Astra 这样的新模型在处理计算机操作、图像理解和复杂任务方面能力突出,但它们给实际的软件工程带来了显著挑战。模型似乎缺乏明确的约束或惩罚机制来抑制低质量代码的生成,因而对长期任务完成的过度关注反而催生出大量不可用或不可维护的产出。
为测试这些能力的极限,一项为期周末的自管理软件工厂实验暴露出令人担忧的模式。模型在被赋予完全自治权以管理上下文和子智能体时,于 35 小时内生成了 75,000 行代码,消耗了约十亿个 tokens(词元)。尽管活动量巨大,产出却几乎毫无价值。代码表现出许多怪异行为:在编辑 C 代码时依赖复杂的 Python 字符串拼接,用 Python 去执行简单的 shell 命令(本可直接调用),以及各种古怪难读、完全违背既有最佳实践的编码风格。
问题的根源似乎在于这些模型的训练目标。它们因词元效率和任务完成而被大量奖励,常采用在临时工具调用中可行但用于持久代码库时灾难性的"代码高尔夫"技巧。这导致可读性明显退化:模型优先选择紧凑、对机器高效的方案,而非便于人类理解的逻辑。在无人看管的情况下,这些智能体会呈现出递归、螺旋式的行为,生成越来越晦涩的任务结构,并将随机常量或任意逻辑混入生产级代码中。
这一趋势对 AI 在软件开发方向上的走向提出了根本性疑问。尽管模型快速演进,但其发展轨迹似乎与以人为中心的工程流程相悖。运行这些智能体的成本——无论是高昂的财务开销,还是需要的大量人类监督——都使其在专业软件工作中的合理性受质疑。看起来这些工具更像是在为其他领域(如 3D rendering 或通用计算自动化)进行优化,反而让专业开发者面临一个越来越难以审计与信任的产品。
归根结底,智能体被优化出来的产物与人类工程师所需之间的脱节在扩大。如果代码最终只成为其他智能体读写的媒介,负责长期维护的人类程序员就可能丧失对其的实用价值。只要这些模型把"完成任务"的原始得分置于连贯性与可维护性之上,它们很可能更多地成为生产"糟粕"的新奇工具,而非严肃软件工程中的可靠伙伴。
The current state of AI engineering is increasingly characterized by what can be described as Neijuan, or involution, where systems demand greater effort and competitive output without any real increase in value or productivity. While new models like GPT 6 Astra are undeniably impressive in their capacity to handle computer use, image understanding, and complex tasks, they present significant challenges for practical software engineering. These models seem to lack clear constraints or penalties for producing low-quality code, leading to an environment where the focus on long-horizon task completion results in the creation of vast amounts of unusable or unmaintainable material.
In an effort to test the limits of these capabilities, a weekend-long experiment running a self-managed software factory revealed concerning patterns. Given full agency to manage context and subagents, the model produced 75,000 lines of code over 35 hours while burning through roughly a billion tokens. Despite this immense activity, the output was effectively valueless. The code displayed strange behaviors, such as a reliance on convoluted Python string splicing for C code editing, unnecessary use of Python to execute simple shell commands, and bizarre, hard-to-read coding styles that deviate entirely from established best practices.
The core issue appears to stem from how these models are trained. They are heavily rewarded for token efficiency and completing tasks, often using "codegolfing" techniques that work well for temporary tool calls but are catastrophic when applied to permanent codebase development. This leads to a degradation of readability, as the models prioritize compact, machine-efficient solutions over human-understandable logic. When left to run unattended, these agents exhibit a recursive, spiraling behavior, creating increasingly obscure task structures and integrating random constants or arbitrary logic into production-level code.
This trend raises fundamental questions about the direction of AI in software development. While these models are evolving rapidly, their trajectory seems at odds with current human-centric engineering processes. The cost of running these agents, both in terms of financial expense and the necessity for intense human oversight, makes them difficult to justify for professional software work. It appears these tools are being optimized for other domains, such as 3D rendering or general computer automation, leaving the professional developer with a product that is increasingly difficult to audit or trust.
Ultimately, there is a growing disconnect between what these agents are optimized to produce and what human engineers require. If code becomes a medium intended solely for other agents to read and write, it may eventually lose all utility for the human programmers responsible for its long-term maintenance. As long as these models prioritize raw completion over coherence and maintainability, they will likely remain more of a curiosity for "slop" production rather than a reliable partner in serious software engineering.
- 在复杂的代码库中,进展常常停滞,因为自动编码代理生成的代码难以维护且质量低下,随着这些代码的累积,重构和调试变得越来越困难。
- 虽然支持者宣称这些工具能提高生产力,但"代理化"方法往往把成本转移到团队里的真实成员身上,产生一种不可避免会崩溃的"新手运气"期,从而积累出难以治理的技术债务。
- 有经验的开发者发现,在缺乏严格人工监督的情况下依赖大型语言模型(LLM)写代码,会削弱开发者对系统的理解,最终拖慢长期开发进度。
- 最新的前沿模型在完成任务时常表现出"盲目热情":它们倾向于生成复杂、难读或冗余的脚本来修改文件,而不是用标准的、可审查的工具。
- 模型训练中的奖励机制似乎更看重长周期任务的完成,而非代码的可读性或可维护性,这使得代理更像"白痴学者(idiot savants)"而非可靠的协作成员。
- 开发工作流已转向密集的提示工程和对代理的"培养",这些工作往往与手动编码一样耗时,但结果却更不确定。
- 关于"AGI"或"革命性"能力的市场宣传常与核心软件工程任务中 token 成本上升、模型性能下降的现实发生冲突。
- 让 AI 代理生成未经人工审查的代码会导致一种"内卷"现象:不断加大的投入与资源消耗并未换来相应的业务价值增长。
- 当这些工具被限定在既有代码库中具体且定义明确的任务时,其战略性价值最大;相反,对新特性进行一次性或完全自主的生成,往往只会产生过度工程化的"淤泥(slop)"。
- 支持者常用的"编译器"类比有其局限:编译器是确定性的、以机器层面的输出为准,而 LLM 生成的是需要开发者维护和理解的、面向人的源代码。
总的来说,这场讨论反映出 AI 编码工具在带来快速初始效用的同时,也在专业软件工程实践中造成越来越多的摩擦。各方虽然普遍承认这些模型在特定任务和探索场景中很有用,但对于它们在无需大量人工投入下,能否产出可维护且高质量的生产级软件仍持深刻怀疑。许多开发者注意到,模型越来越"古怪"、难以协作,它们更注重完成任务而非维护架构完整性和透明度。归根结底,行业对耗费大量 token 的代理式工作流的狂热,与对可读性和可维护性忽视的现实形成了悖论:工程组织不得不投入更多资源来管理日益复杂且脆弱的自动化产出。
• Progress in complex codebases often stalls because automated coding agents generate unmaintainable, low-quality code that becomes increasingly difficult to refactor or debug as it accumulates.
• While proponents claim these tools boost productivity, the "agentic" approach often externalizes costs to human team members, creating a "beginner's luck" phase that inevitably collapses into unmanageable technical debt.
• Experienced developers find that relying on LLMs to write code without strict human oversight results in a degradation of the developer's own understanding of the system, ultimately slowing down long-term development.
• The latest frontier models exhibit a "maniacal" approach to task completion, often preferring to write complex, unreadable, or redundant scripts to modify files rather than using standard, reviewable tools.
• Reward structures in model training appear to prioritize successful long-horizon task completion over code readability or maintainability, leading to agents that act as "idiot savants" rather than collaborative team players.
• Development workflows have shifted toward intensive prompt engineering and "grooming" of agents, which is often as time-consuming as writing the code manually, yet produces less deterministic results.
• Marketing claims of "AGI" or "revolutionary" capability frequently clash with the reality of increasing token costs and declining model performance for core software engineering tasks.
• The practice of using AI agents to generate code that humans do not review is linked to "involution" (or Neijuan), where intensified effort and resource consumption fail to yield a corresponding increase in actual business value.
• Strategic use of these tools is most effective when constrained to specific, well-defined tasks within established codebases, whereas "one-shot" or fully autonomous generation for novel features often results in over-engineered "slop."
• The compiler analogy frequently used by proponents is flawed because compilers are deterministic and prioritize the machine-level output, whereas LLMs generate human-readable source code that must be maintained and understood by developers.
The conversation reflects a growing tension between the rapid, initial utility of AI coding tools and the long-term friction they introduce into professional software engineering. While there is a consensus that these models are powerful enough to be useful for specific tasks and discovery, deep skepticism persists regarding their ability to produce maintainable, high-quality production software without significant human effort. Many developers observe that the models have become increasingly "weird" and difficult to work with, prioritizing the completion of a task over architectural integrity and transparency. Ultimately, the discussion suggests that while the industry is currently obsessed with token-heavy, agentic workflows, the lack of emphasis on readability and maintainability is leading to a paradox where engineering organizations spend more resources to manage increasingly complex and fragile automated outputs.
16 岁的 Ángela Karime Venegas Hernández 是 Altamira, Tamaulipas 地区 CETIS 78 的一名学生,她开发出一种依靠声波而非传统化学品或水的创新灭火装置。她的灵感来自一个简单的好奇心:想在不吹气或不使用常规物质的情况下熄灭蜡烛。 Sixteen-year-old Ángela Karime Venegas Hernández, a student at CETIS 78 in Altamira, Tamaulipas, has developed an innovative fire extinguishing device that relies on sound waves rather than traditional chemicals or water. Her inspiration began with a simple curiosity: the desire to extinguish a candle without blowing on it or using conventional substances.
16 岁的 Ángela Karime Venegas Hernández 是 Altamira, Tamaulipas 地区 CETIS 78 的一名学生,她开发出一种依靠声波而非传统化学品或水的创新灭火装置。她的灵感来自一个简单的好奇心:想在不吹气或不使用常规物质的情况下熄灭蜡烛。
为此,她设计了由 12 伏电池供电的装置,采用频率发生器和经过校准的扬声器,以每秒 30 次脉冲振动工作。该振动能有效将氧气从火焰周围推开,使火焰因缺氧在 5 到 8 秒内熄灭。
她将项目命名为 Vortex Tech,并通过严格测试验证了其多用途性。 Ángela 进行了 100 多次试验,确认装置对多种火灾类型均有效,包括木材、易燃液体、厨房油脂以及对电子设备敏感的火情。
除了速度和效果,这项发明还具有显著的环保和安全优势:通过声波灭火不会造成污染、不留化学残留,也不会对操作人员造成身体伤害。
这个起于学校的小实验如今已受到国际关注。凭借这一突破,Ángela 将代表 Mexico 参加国际科学博览会,展示她应对人类最古老且最持久危险之一的方法。
Sixteen-year-old Ángela Karime Venegas Hernández, a student at CETIS 78 in Altamira, Tamaulipas, has developed an innovative fire extinguishing device that relies on sound waves rather than traditional chemicals or water. Her inspiration began with a simple curiosity: the desire to extinguish a candle without blowing on it or using conventional substances.
To achieve this, she engineered a device powered by a 12-volt battery, utilizing a frequency generator and a speaker calibrated to emit 30 pulses per second. This specific vibration creates an effect that effectively pushes oxygen away from the flame. By depriving the fire of the oxygen it needs to burn, the device is able to extinguish flames in just five to eight seconds.
The project, which she has named Vortex Tech, has proven to be remarkably versatile through rigorous testing. Ángela conducted over 100 trials to confirm the device's efficacy across various types of fires, including those involving wood, flammable liquids, cooking grease, and sensitive electronic equipment.
Beyond its speed and effectiveness, the invention offers significant environmental and safety benefits. Because it operates through sound waves, the process does not create pollution, leave behind chemical residues, or pose any physical harm to the person operating the device.
What started as a modest school experiment has now gained international attention. As a result of her breakthrough, Ángela is set to represent Mexico at an international science fair, where she will showcase her method for combating one of humanity's oldest and most persistent dangers.
"声波灭火器"这一概念已有充分记录,早在十多年前就有研究和演示,参与者包括 DARPA 、大学生团队和 MythBusters team 等。该技术利用低频声波破坏火焰与燃料之间的边界层,通过物理作用将氧气排开以抑制燃烧。但它存在明显的现实局限:声波本身并不提供冷却作用,如果没有降低燃料温度的辅助手段,一旦声脉冲停止,火焰通常会立即复燃。
尽管声学装置在像实验室这样的受控环境中可能有一定的利基应用(类似于 CO2 灭火器的用途),但在大规模、现实世界的消防场景中,目前并不实用。媒体往往夸大这些学生项目,把它们包装成新奇的"发明",这让那些了解其物理原理及其悠久历史的人感到无奈和沮丧。学生们的独立发现具有显著的教育价值——主要在于科学探索和问题解决的过程,而不是创造出可商业化或独一无二的技术突破。
对这些报道的批评通常指向新闻写作中专业性不足或夸大其词,而不是否定年轻发明者本身的努力。讨论经常演变成关于人口统计刻板印象和智商指标的非建设性争论,凸显出一种反复出现的模式:令人欣慰的科学故事容易引发关于国家或个人价值的极化反应。高质量的科学传播难以实现,需要在吸引人的实验展示与严谨、清晰且适合年轻受众的解释之间找到平衡。
此外,AI-generated content 的泛滥以及网站上激进的"human verification"门槛,为在线讨论增加了一层摩擦和不信任,常常分散人们对实际问题的注意力。围绕这项发明的讨论反映了鼓励青年科学好奇心与技术现实之间的经典张力:大多数人认为学生的实验是值得赞赏的教育活动,但媒体将其误读为独特或革命性发现是不恰当的。尽管声波抑制的物理原理成立,但缺乏防止复燃所需的冷却手段仍是一个重大工程障碍。最终,这些讨论也对现代媒体周期提出了批评:为了迎合感人的叙事,媒体常常牺牲技术准确性,从而让熟悉该技术历史的观察者感到疲惫。
• Acoustic fire extinguishers are a well-documented concept, with prior research and demonstrations conducted by DARPA, university students, and the MythBusters team dating back over a decade.
• The technology operates by using low-frequency sound waves to disrupt the boundary layer between the flame and its fuel source, physically pushing oxygen away to suppress combustion.
• Significant practical limitations exist because sound suppression does not provide cooling; without a secondary method to lower the fuel temperature, fires frequently reignite immediately once the sonic pulses cease.
• While acoustic devices may have niche applications for delicate environments like laboratories—similar to CO2 extinguishers—they are currently ineffective for large-scale, real-world fire suppression.
• Public discourse often frames these student projects as novel "inventions" due to media exaggeration, leading to skepticism and frustration from those who recognize the long history of the underlying physics.
• Independent discovery by students holds educational value, as the primary benefit is the process of scientific exploration and problem-solving, rather than the expectation of creating a commercialized or unique breakthrough.
• Criticism directed at these reports generally targets the journalistic incompetence or hyperbole in the writing rather than the efforts of the young inventors themselves.
• Discourse frequently shifts into unproductive debates over demographic stereotypes and IQ metrics, highlighting a recurring pattern where feel-good science stories trigger polarized reactions regarding national or individual merit.
• High-quality educational content in science communication is difficult to achieve, requiring a balance between entertaining experiments and rigorous, clear explanations that remain accessible to young audiences.
• The increasing prevalence of AI-generated content and aggressive "human verification" gatekeeping by websites adds a layer of friction and distrust to online discussions, often distracting from the actual subject matter.
The conversation surrounding this invention reflects a classic tension between the encouragement of youthful scientific curiosity and the technical reality of innovation. While many participants acknowledge that the student's experiment is a valid and commendable educational endeavor, there is a clear consensus that the media coverage misrepresented the work as a unique or revolutionary discovery. The discussion emphasizes that while the physics of acoustic suppression is sound, the inability of sound waves to provide the necessary cooling required to prevent reignition remains a major engineering hurdle. Ultimately, the thread serves as a critique of modern media cycles that prioritize "feel-good" narratives over technical accuracy, leading to predictable fatigue among those familiar with previous iterations of the technology.
Google 宣布在 Europe 的有史以来最大单笔投资,承诺投入 130 亿欧元用于加强 Finland 的 AI 基础设施。该投资包括建设三座新数据中心、扩建位于 Hamina 的现有设施,并支持多项清洁能源项目,旨在满足对 Gemini 聊天机器人、 Google Search 、 Maps 和 YouTube 等 AI 驱动服务日益增长的全球需求。 Google has unveiled its largest single investment in Europe, committing 13 billion euros to boost AI infrastructure in Finland. The comprehensive investment package includes the construction of three new data centers, the expansion of an existing facility in Hamina, and support for various clean energy projects. This move is designed to meet the skyrocketing global demand for AI-driven services, such as the Gemini chatbot, Google Search, Maps, and YouTube.
Google 宣布在 Europe 的有史以来最大单笔投资,承诺投入 130 亿欧元用于加强 Finland 的 AI 基础设施。该投资包括建设三座新数据中心、扩建位于 Hamina 的现有设施,并支持多项清洁能源项目,旨在满足对 Gemini 聊天机器人、 Google Search 、 Maps 和 YouTube 等 AI 驱动服务日益增长的全球需求。
计划的关键一环是与公用事业公司 Fortum 签订的一项为期 22 年的长期合约。根据协议,Google 将购买 Loviisa 核电站多达 50% 的发电量。这笔交易不仅为 Google 的高耗能业务提供了稳定电力,也为 Fortum 提供了增加发电能力并延长电厂运营寿命所需的财务保障。
Finland 正逐渐成为大型科技公司青睐的数据中心选址。 Prime Minister Petteri Orpo 强调了该国的优势:凉爽气候有助于自然降低大规模服务器的冷却成本,可靠且低碳的电力供应和稳定的数字基础设施也极具吸引力。最近 TikTok 也在当地投资了 10 亿美元。
Google 预计该项目将带来显著的经济效益,在 2027 年和 2028 年的施工期预计将创造超过 37,000 个就业机会。除了实体基础设施外,投资还将设立社区基金,用于支持当地生物多样性、科研、教育与劳动力培养。这一做法契合 Alphabet 更广泛的全球战略,即在竞争中大幅增加投入,以建立人工智能革命所需的产能。
Google has unveiled its largest single investment in Europe, committing 13 billion euros to boost AI infrastructure in Finland. The comprehensive investment package includes the construction of three new data centers, the expansion of an existing facility in Hamina, and support for various clean energy projects. This move is designed to meet the skyrocketing global demand for AI-driven services, such as the Gemini chatbot, Google Search, Maps, and YouTube.
A critical component of this initiative is a long-term, 22-year partnership with the Finnish utility firm Fortum. Under this agreement, Google will purchase up to 50 percent of the energy produced by the Loviisa nuclear power plant. This deal not only secures a stable power supply for Google's energy-intensive operations but also provides the utility company with the financial certainty needed to increase the plant's generating capacity and extend its operational life.
Finland has become an increasingly attractive destination for major technology companies looking to house their data centers. Prime Minister Petteri Orpo highlighted the country's strengths, including its cool climate, which naturally helps reduce the cooling costs associated with massive server arrays. Additionally, the nation offers a reliable, low-carbon electricity supply and a stable digital infrastructure, which recently also drew a 1 billion dollar investment from TikTok.
The economic impact of Google's project is expected to be significant, with the company anticipating the creation of over 37,000 jobs during the construction phases in 2027 and 2028. Beyond physical infrastructure, the investment also allocates resources for community funds that will support local biodiversity, research, education, and workforce development. This approach aligns with Alphabet's broader global strategy to ramp up spending significantly as it competes to build the necessary capacity for the ongoing artificial intelligence revolution.
• Google 与 Fortum 的合作旨在购买 Loviisa 核电厂的电力输出,为公用事业公司提供资助厂站寿命延长项目所需的资金确定性,从而有效确保该设施在 2030 年后继续运营。
• Finland 因寒冷气候有助于降低冷却成本且电网碳强度较低,越来越被视为数据中心的理想选址;但这种集中对国内电价和电网稳定性的长期影响仍存在争议。
• 围绕此类交易究竟是真正增加了清洁能源产能,还是仅仅替代了现有的核能需求存在分歧——后者可能会在用电高峰时段迫使其他电力用户转而依赖碳排放更高的能源。
• 批评者认为,尽管数据中心是资本密集型项目,但相比传统制造业,其带来的长期就业机会有限,因此质疑主办国能否从中获得足以抵消对电力基础设施潜在压力的经济利益。
• 欧洲的碳排放总量控制与交易(cap-and-trade)体系被视为限制数据中心推动高污染发电量上升的一道防线,因为无论电力被哪个行业消费,二氧化碳排放总量都有严格上限。
• 核能的经济可行性经常受到质疑:建设成本高、准备周期长且监管复杂,因此有人认为风能、太阳能加上电池储能在应对现代能源需求方面更快且更具成本效益。
• 核能支持者则强调其提供必要的基荷容量,而间歇性可再生能源——即便配备电池储能——在季节性波动(尤其是北部气候条件下)期间仍难以稳定替代这种容量。
• 对数据中心的公众反对常被视为对公共资源"corporate capture"的反应,即科技巨头获得能源使用优先权,而当地社区却要承受电价波动或电网限制的风险。
• 关于环境与安全的争论两极分化:一方面有人以 Chernobyl 等历史性事故证明核能的内在危险;另一方面也有人引用死亡率统计,指出煤炭和其他化石燃料因污染导致的死亡更多。
• 由于与 Russia 的持续紧张关系,Nordic 地区对能源依赖的地缘政治担忧加剧,促使各国更倾向于发展能够增强能源主权与安全的基础设施。
这场讨论反映出两个阵营的尖锐分歧:一方认为 AI 驱动的数据中心投资可成为推动关键能源基础设施升级的催化剂,另一方则担心这些设施会消耗公共资源并推高国内用户的能源成本。尽管各方都认为 Finland 的电网相当清洁,但对于这些企业合同究竟是在实质上支持向可持续能源的转型,还是仅为科技公司提供公关利益,仍存在严重分歧。归根结底,这场争论凸显了在计算需求不断增长的背景下,技术、经济与地缘政治如何交织影响每一项重大基础设施决策。
• Google's partnership with Fortum to purchase output from the Loviisa nuclear plant provides the utility with the financial certainty needed to fund multi-year life-extension projects, effectively ensuring the facility remains operational beyond 2030.
• Finland is increasingly viewed as an ideal location for data centers due to its cold climate, which lowers cooling costs, and its low-carbon energy grid, though the long-term impact on domestic electricity prices and grid stability remains a subject of debate.
• There is a disagreement regarding whether such deals genuinely increase clean energy capacity or merely displace existing nuclear demand, potentially forcing other grid users to rely on more carbon-intensive sources during periods of high demand.
• Critics argue that data centers, while capital-intensive, offer limited long-term employment compared to traditional manufacturing, questioning whether the economic benefits for the host country justify the potential strain on electricity infrastructure.
• The European "cap-and-trade" system for carbon emissions is cited as a mechanism that prevents data centers from driving overall increases in dirty energy production, as total CO2 output remains strictly limited regardless of which sector consumes the power.
• The economic feasibility of nuclear power is frequently challenged by its high construction costs, long lead times, and regulatory complexities, with some arguing that wind, solar, and battery storage are faster and more cost-effective solutions for modern energy needs.
• Proponents of nuclear power emphasize that it provides essential baseload capacity that intermittent renewables—even with battery storage—struggle to match reliably during seasonal fluctuations, especially in northern climates.
• Public opposition to data centers is often framed as a reaction to perceived "corporate capture" of public resources, where tech giants secure preferential access to energy while local communities face price volatility or grid constraints.
• Arguments regarding environmental safety are polarized, with some highlighting historical nuclear accidents like Chernobyl as proof of inherent danger, while others point to mortality statistics showing coal and other fossil fuels cause significantly more deaths through pollution.
• Geopolitical concerns regarding energy dependence are heightened in the Nordic region due to the ongoing tensions with Russia, leading to a strong national preference for infrastructure that reinforces energy sovereignty and security.
The discussion reflects a sharp divide between those who view AI-driven data center investments as a catalyst for essential energy infrastructure upgrades and those who fear these facilities will strain public resources and inflate energy costs for domestic users. While there is a consensus that Finland's grid is exceptionally clean, participants strongly disagree on whether these corporate contracts effectively support a transition to sustainable energy or merely offer a public relations benefit for tech companies. Ultimately, the conversation highlights the broader challenges of modernizing energy grids in the face of rising computational demand, with technical, economic, and geopolitical concerns intertwined in every major infrastructure decision.
尝试用本地 large language model 替代编码工具框架中对远程 API 的调用,往往会令人很沮丧。虽然像 llama-bench 这样的基准测试可能显示很高的 tokens-per-second,但真实的使用体验常常被明显的延迟、卡顿和漫长的等待时间破坏。这种差距主要因为大多数编码工具框架在设计时假定会运行在高速的数据中心 GPU 上,而非受限的消费级笔记本上。 Attempting to use a local large language model as a replacement for remote API calls in coding harnesses often leads to frustration. While benchmarking tools like llama-bench might suggest high tokens-per-second, the actual user experience is frequently marred by significant latency, stalling, and long wait times. This discrepancy exists primarily because most coding harnesses were built with the assumption of high-speed data center GPUs, rather than the constraints of a consumer laptop.
尝试用本地 large language model 替代编码工具框架中对远程 API 的调用,往往会令人很沮丧。虽然像 llama-bench 这样的基准测试可能显示很高的 tokens-per-second,但真实的使用体验常常被明显的延迟、卡顿和漫长的等待时间破坏。这种差距主要因为大多数编码工具框架在设计时假定会运行在高速的数据中心 GPU 上,而非受限的消费级笔记本上。
本地硬件的物理特性会把系统逼到三种典型的设计困境上。第一,巨大的系统提示(system prompt)和复杂的工具 schema 带来沉重的预填充负担。当笔记本处理 token 的速度只是数据中心的一小部分时,一个 18,000 token 的提示可能导致模型在开始生成代码前就要等待几分钟。第二,这类大提示会占用大量上下文窗口,几乎没剩多少空间用于实际任务。第三,许多框架会不断发起会话标题或摘要的侧请求,在本地运行时这些请求会把模型推入频繁且重叠的队列,导致重复的长时间预填充,使得性能降到不可用的程度。
为了解这些瓶颈,在一台高端 M4 MacBook Pro 上,使用量化的 Qwen 3.8 27B 模型对九个流行的编码框架进行了多项 Python 练习测试。结果把这些工具分为三类:精简且稳定的框架保持短小的系统提示并高度重用缓存,是本地使用的最佳候选;尽管提示较长或工具 schema 较多但有纪律的重型工具,在完成初次预填充后仍能正常工作;而启动沉重型的工具可能需要三到四分钟的静默等待,代理才能迈出第一步,因而不适合本地工作流。
问题的核心不是工程水平低下,而是这些工具最初的设计环境不同。当延迟以毫秒计时,几乎免费的特性在本地硅片上以秒或分钟计时就成了沉重的负担。因此,像 chad 这样的项目应运而生,采取稀缺性思维,限制工具表面并针对持久化的前缀缓存(prefix caching)进行优化。这类设计把效率放在首位,确保开发循环保持响应,而不是被模型的开销阻塞。
真正的性能突破发生在框架与推理引擎运行在同一进程时。通过打破传统的客户端 - 服务器模型,将引擎内嵌到进程中(例如通过 MLX),可以将与请求模式相关的延迟降到最低。那些拥有自身缓存并采用专门草稿生成技术的框架,能够实现显著更高的 tokens-per-second,把曾经缓慢、令人沮丧的体验变成流畅高效的编码环境。
Attempting to use a local large language model as a replacement for remote API calls in coding harnesses often leads to frustration. While benchmarking tools like llama-bench might suggest high tokens-per-second, the actual user experience is frequently marred by significant latency, stalling, and long wait times. This discrepancy exists primarily because most coding harnesses were built with the assumption of high-speed data center GPUs, rather than the constraints of a consumer laptop.
The physics of local hardware often pits the system against three specific design choices. First, large system prompts and extensive tool schemas create a massive prefill burden. When a laptop processes tokens at a fraction of the speed of a data center, a 18,000-token prompt can result in several minutes of waiting before the model even begins to generate code. Second, these large prompts consume a significant portion of the available context window, leaving little room for actual task execution. Third, many harnesses are designed to make constant side requests for session titles or summaries, which, when run locally, force the model into frequent, overlapping queues and repeated long prefills that degrade performance to the point of being unusable.
To better understand these bottlenecks, nine popular coding harnesses were tested against a series of Python exercises on a high-end M4 MacBook Pro using a quantized Qwen 3.8 27B model. The results categorized these tools into three distinct groups. Lean and stable harnesses maintain trim system prompts and high cache reuse, making them the most viable candidates for local use. Heavy but disciplined tools, while burdened by long prompts or excessive tool schemas, remain functional once the initial prefill is complete. Finally, heavy-to-start tools can require three to four minutes of inactivity before the agent takes its first step, rendering them impractical for a local workflow.
The core issue is not poor engineering, but rather the environment for which these tools were originally designed. Features that are virtually free when latency is measured in milliseconds become heavy liabilities when measured in seconds or minutes on local silicon. Consequently, projects like chad have emerged to embrace a scarcity mindset, limiting tool surfaces and optimizing for persistent prefix caching. These designs prioritize efficiency to ensure that the development loop remains responsive rather than blocked by the model's overhead.
The true breakthrough in performance, however, arises when the harness and the inference engine share the same process. By breaking away from the standard client-server model and moving the engine in-process, such as through MLX, the latency associated with request patterns is minimized. Harnesses that own their cache and utilize specialized drafting techniques can achieve significantly higher token-per-second rates, turning what was once a sluggish, frustrating experience into a fluid and efficient coding environment.
轻量级的编码 agents 和基于终端的 TUI harnesses 在资源受限的环境(如笔记本电脑和 SBC)中越来越流行,像 hax 和 clm 这样的项目提供极小的可执行文件(小于 1MB)且内存占用极低。
开发者正从臃肿且依赖繁重的工具转向"表现良好"的 Unix 风格实用程序,这类工具遵循 XDG 标准并避免使用侵入式安装脚本。
基于 FUSE 的 overlay 虚拟文件系统正成为提高 agent 安全性的突破性功能,允许用户"回滚"会话或防止误删,且无需复杂的容器化或手动的 git 检点。
Token 效率是主要争论点,许多自称"功能齐全"的 harnesses 因开销过高而受到批评,用户反映"精简"配置在实际任务中更具成本效益且性能更好。
系统提示(system prompts)往往是导致臃肿和幻觉(hallucination)的根源,一些用户倾向于将这些提示模块化为多个文件,以便更好地控制并在不同 agent 系统间轻松迁移。
Qwen 27B 和 35B 模型因其性能与计算开销比在本地推理中受到青睐,通常被用作可以在消费级硬件上运行的"美化版 linter"或编码助手。
社区明确希望有一个标准化且可复现的基准套件,用来衡量各类 harness 的每轮前缀 token 数、首个 token 响应时间以及缓存重用情况。
开发者之间存在明显分歧:一部分人优先"vibe coding"(快速迭代和试验),另一部分人则反感 AI 生成的文档或"马虎"的内容,认为由人工撰写的内容是项目可信度的基准。
网站质量,特别是 Notion 托管页面,经常被指出为主要摩擦点,用户抱怨滚动体验差、渲染卡顿和导航失效。
Agent harnesses 市场充斥着许多仅在界面上略有差别的定制工具,真正的创新来自那些从根本上改变 agent 与操作系统交互方式或状态处理方式的项目。
目前 AI 编码 agents 的格局以轻量、基于终端的工具激增为特征,原因是开发者对大型框架的臃肿和性能问题感到厌倦。业界普遍认为"vibe coding"是一种有效的开发方法论,但它常被不一致的基准测试、糟糕的 token 管理以及工具中不可靠的用户体验所制约。尽管开发者在尝试用虚拟化文件系统等高级技术提升安全性和工作流,但他们对增加可能成为维护负担的复杂性仍保持谨慎。总体上,社区倾向于精简且高度可控的实用工具,支持模块化定制,这反映出对一刀切平台的抵制。
• Lightweight coding agents and TUI harnesses are increasingly popular for resource-constrained environments like laptops and SBCs, with projects like hax and clm providing minimal binaries (sub-1MB) with low memory footprints.
• Developers are shifting away from bloated, dependency-heavy tools in favor of "well-behaved" Unix-style utilities that respect XDG standards and avoid intrusive install scripts.
• Virtualized filesystems using FUSE-based overlays are emerging as a game-changing feature for agent safety, allowing users to "rewind" sessions or prevent accidental deletions without needing complex containerization or manual git checkpoints.
• Token efficiency is a major point of contention; many "full-fledged" harnesses are criticized for excessive overhead, with users reporting that "barebones" setups are significantly more cost-effective and performant for real-world tasks.
• System prompts are often a source of bloat and hallucination; some users prefer modularizing these prompts into multiple files for better control and easier migration between different agent systems.
• Qwen 27B and 35B models are currently favored for local inference due to their performance-to-compute ratio, often being used as "glorified linters" or coding assistants that can run on consumer-grade hardware.
• There is a expressed desire for a standardized, reproducible benchmark suite that can measure per-turn prefix token counts, time-to-first-token, and cache reuse across various harnesses.
• A clear divide exists between developers prioritizing "vibe coding" (quick iteration and experimentation) and those who are alienated by AI-generated documentation or "slop," viewing human-authored content as a baseline for project credibility.
• Website quality, particularly the use of Notion-hosted pages, is frequently cited as a major friction point, with users complaining about poor scrolling, janky rendering, and broken navigation.
• The market for agent harnesses is saturated with many custom tools that differ only in UI; genuine innovation is seen in projects that fundamentally change how the agent interacts with the OS or handles state.
The landscape for AI coding agents is currently characterized by a proliferation of lightweight, terminal-based tools as developers grow frustrated with the bloat and performance issues of larger frameworks. There is a strong consensus that "vibe coding" is a valid development methodology, but it is often hampered by inconsistent benchmarking, poor token management, and unreliable UX in the tools provided. While developers are experimenting with advanced techniques like virtualized filesystems to enhance safety and workflow, they remain cautious about adding complexity that could become a maintenance burden. Ultimately, the community is moving toward a preference for lean, highly controllable utilities that allow for modular customization, signaling a rejection of one-size-fits-all platforms.
236 comments • Comments Link
• 关于编程已经"solved"的说法受到广泛质疑:目前的 AI 模型在高层架构决策、长期可维护性以及将复杂功能集成到大型既有系统方面仍然力不从心。
• 虽然 AI 能有效完成一次性任务和简单脚本,但在缺乏监督时常会产出冗余、低质量的代码,进而增加技术债务并使对最终软件进行逻辑推理变得更困难。
• 行业内缺乏衡量代码质量的稳健客观指标;单靠代码行数(LOC)等简单指标会触发 Goodhart's Law,导致模型倾向于优化指标而非真正的质量。
• 在软件开发中有效使用 AI 需要积极的人为监督,人在流程中应扮演导演或架构师的角色,而不是单纯的打字员,这通常要求高度自律与辅助工具来维护代码完整性。
• 编程在团队沟通中是一个有损通道,随着 AI 接管越来越多的实现工作,维持人类团队对系统的连贯认知变得愈发关键且艰难。
• 企业级软件开发受制于遗留系统、业务需求和高可靠性要求等复杂约束,目前的 agents 通常无法自主管理这些需求。
• AI 编码代理在"在正确的地方进行修改"方面表现欠佳,常常违背既定架构模式,因为它们缺乏对全局设计和原始系统意图的深入理解。
• 对 AI 的看法存在显著分歧:一方把它视为变革性的生产力工具,认为通过提高可能性的下限从而从根本上"solved"了编程;另一方则强调,软件工程的核心挑战——可靠性、安全性与长期维护——尚未解决,仍需深厚的人类专业知识。
• 一些开发者认为人类编写的企业级代码历来质量堪忧,暗示尽管 AI 目前有缺陷,最终可能提升平均水准。
• 当前的 AI 热潮导致代码量呈指数级增长,但维护和调试 AI 生成系统的长期成本正在显现,可能掩盖短期开发速度带来的收益。
关于编程是否已被"solved"的争论,本质上源于对"软件开发"定义的根本分歧。一种观点把编程看作产出功能性结果的行为,从这个角度来看,AI 能生成可运行(尽管有时臃肿)的代码是一场决定性胜利。另一种观点则认为编程只是软件工程的一部分,软件工程还包括为可维护性而设计、保证可靠性以及应对复杂的组织与业务需求等更艰巨的任务。 AI 对熟练开发者是强大的倍增器,但业界普遍认为它无法替代人类在架构与问题解决上所需的全面理解,尤其是在关键的生产环境中。这场讨论反映了 AI 驱动开发带来的即时速度与传统以人为本的精准性、设计连贯性及长期系统健康价值之间的紧张关系。 • The premise that coding is "solved" is widely contested, as current AI models struggle with high-level architectural decisions, long-term maintainability, and the complex integration of features into large, existing systems.
• While AI effectively handles one-shot tasks and simple scripts, it often produces "slop" or verbose, low-quality code when left unsupervised, which can lead to increased technical debt and difficulty in reasoning about the resulting software.
• The industry currently lacks robust, objective metrics for code quality; relying on simple indicators like lines of code (LOC) often triggers Goodhart's Law, where models optimize for the metric rather than genuine quality.
• Effective AI usage in software development requires active human oversight, where the human acts as a director or architect rather than a mere typist, often necessitating significant discipline and secondary tooling to maintain code integrity.
• Coding is a lossy channel for team communication; as AI takes over more of the implementation, maintaining a coherent mental model of the system becomes a critical, yet increasingly difficult, challenge for human teams.
• Enterprise-level software development involves complex constraints like legacy systems, business requirements, and high-stakes reliability needs, which current agents are generally not equipped to manage autonomously.
• AI coding agents struggle with "making changes in the right places," often working against established architectural patterns because they lack a full understanding of the global design and the underlying intentions of the original system.
• There is a notable divide between those who view AI as a transformative productivity tool that has fundamentally "solved" coding by raising the floor of what is possible, and those who emphasize that the core challenges of software engineering—reliability, security, and long-term maintenance—remain unsolved and require deep human expertise.
• Some developers argue that human-written enterprise code has historically been of poor quality, suggesting that AI might eventually improve the average standard despite its current shortcomings.
• The current AI boom has led to an exponential increase in the volume of code, but the long-term cost of maintaining and debugging AI-generated systems is an emerging concern that may overshadow the short-term gains in development speed.
The debate over whether coding is "solved" hinges on a fundamental disagreement regarding the definition of software development. One perspective views coding as the act of producing functional output, where AI's ability to generate working, albeit sometimes bloated, code represents a decisive victory. Conversely, others argue that coding is merely a subset of software engineering, which encompasses the much harder tasks of designing for maintainability, ensuring reliability, and navigating complex organizational and business requirements. While AI acts as a powerful multiplier for skilled developers, the consensus among practitioners is that it lacks the holistic understanding necessary to replace the human role in architecture and problem-solving, particularly in critical production environments. The discussion reflects a tension between the immediate velocity of AI-driven development and the traditional, human-centered values of precision, design coherence, and long-term system health.