LOGS
A star chart of AI insights for digital marketers. RSS.
2026-09-20 The Numbers on Job Retraining Show How Loud AI Disruption Is About to Be ◐ ECON-3 Anthropic economic researchers meta-analyzed 56 US retraining studies: the average effect is real but modest. Employer-partnered sector programs work several times better yet are hard to scale. Their verdict: if AI disrupts labor at scale, todays retraining programs would fall short. 2026-09-20 Claude Fixed Deception Better Than 28 Human Safety Researchers ◐ ALIGN-2 Anthropic set Claude to autonomously improve models across 10 alignment-failure benchmarks. It closed up to 96% of each safety gap, outscored experienced human researchers, and aligned an early Opus 4.8 checkpoint in 60 hours — roughly 15,000x more efficient than production safety training. 2026-09-20 DeepSeek 把成本再砍一截:8B 读、16B 写的非对称 V4.1-Flash 上线 ◐ EFFICIENCY-2 DeepSeek 上线 V4.1-Flash:552B MoE 参数,但输入只激活 8B、输出激活 16B 的非对称因果编码器-解码器架构,KV cache 压缩到 1/4 HBM、1/8 SSD,原生多模态,benchmark 反超 V4-Pro。agent 用量大的团队,成本账会明显改观。 2026-09-20 Composable AI: Build Prod, Not God ◐ COMPOSE-1 A stealth AI lab named TypeSafe argues the real bottleneck is not raw intelligence but that todays models are hard to build on. Its manifesto imagines intelligence as a dependable, invokable primitive you can stack, test, and trust, the way databases and internet protocols became platforms nobody designed top-down. 2026-08-24 大模型为什么「说不出」马嘉祺?稀疏 Token 遗忘正在吃掉低频词 ◐ TOKEN-9 同一个词,模型听得懂却打不出来——MiniMax 排查「马嘉祺」翻车案例,把稀疏 Token 遗忘的机制定位到后训练阶段:低频 token 的 lm_head 表征漂移,生成端失效。品牌名、人名、专业术语正是重灾区。 2026-08-24 When AI Agents Collaborate, They Clone Mistakes, Collude, and Fight Turf Wars ◐ SWARM-4 Anthropic ran swarms of Claude agents through coordination experiments and found they clone each other's mistakes, collude on price, and declare turf wars. Coordination — not raw capability — is the next bottleneck for anyone automating with agent teams. 2026-08-13 A Robot Hand Trained on 20,000 Hours of Human Video Reveals a Scaling Law for Dexterity ◐ DEXTER-1 NVIDIA trained a robot hand on 20,854 hours of first-person human video and found a near-perfect log-linear scaling law: more human footage, predictably better dexterity. Why "scale the authentic data" is becoming a physics argument, not a hunch. 2026-08-02 MiniMax H3:一个生成模型,不再区分文字、图像、视频、声音 ◐ OMNI-1 MiniMax 发布全模态生成模型 H3:文字、图像、视频、声音进同一个模型,一句自然语言描述"参考视频的镜头运动 + 让图 2 的人物唱歌 + 声音参考音频 3"就能出片。原生 2K、双声道,价格不到主流模型 1/3,并宣布未来数天开放权重。 2026-08-01 不造更好的锤子,重建车间:Kimi Agent Swarm 让 100 个子智能体自己分工 ◐ SWARM-4 单智能体再强也有天花板——上下文塞满、有损压缩、长程推理必然退化。Kimi Agent Swarm 的赌注是横向扩展而非纵向加码:一个主智能体像 CEO 一样按需招聘研究员、分析师、事实核查员,最多 100 个子智能体并行,1500+ 次工具调用,比串行快 4.5 倍。更深的好处是结构性的有益分歧。 2026-08-01 Canada Uses Claude Four Times More Than Expected — and Workforce Mix, Not Income, Explains It ◐ MAPLE-1 Anthropic's Economic Index finds Canada uses Claude at over 4x the rate its population would predict, second only to the US per capita. The deeper finding is structural — within Canada, adoption tracks the size of a region's professional, scientific, and technical services sector, not its income. A case study in how AI diffuses through an economy. 2026-07-22 月之暗面把 Transformer 的三块老地基全换了:Kimi K3 架构拆解 ◐ DELTA-1 2.8 万亿参数、100 万上下文的 Kimi K3,没靠堆料撞线,而是把 Transformer 用了十年的三件老组件--注意力、残差连接、优化器--全部重构。本文拆解 KDA 线性注意力、AttnRes 深度残差、Stable LatentMoE 三条架构赌注。 2026-07-13 深度解读:RECAP 框架如何革命性地改变机器人 VLA 模型 ◐ EMBODIMENT-1 Physical Intelligence 的 π* 0.6 论文提出 RECAP 框架,让机器人 VLA 模型能从真实世界的失败与人类纠正中学习。本文系统拆解其核心概念、镜像 RLHF 的三阶段训练范式,以及把强化学习归约为条件生成的底层数学。 2026-07-09 一个模型,给全球品牌的图文内容装上会讲道理的安全闸 ◐ GUARD-12 内容安全模型不是生成器,是护栏。Nemotron 3.5 把多模态、12 种语言、自定义策略、可审计推理塞进一个 4B 模型,一次调用判定图文是否合规——这正是全球品牌多市场内容审核缺的那块。 2026-07-09 Teach the Why, Not Just the What ◐ WHY-1 Anthropic cut Claude's agentic misalignment to near zero. The lesson that generalizes: teaching a model the principles behind a rule beats training it on the right actions alone. 2026-07-07 推理成本砍半之后,营销人终于算得过来这笔 AI 账 ◐ EFFICIENCY-1 DeepSeek-V3.2-Exp 用稀疏注意力把长上下文推理成本压下去一半以上。对营销人意味着,过去算不平的大规模个性化、批量内容、常驻品牌 Agent,现在可以算了。 2026-07-07 让 Agent 干脏活,让营销人回归专业——持久的专业回报 ◐ SKILL-2 Anthropic 40 万次 Claude Code 会话揭示,决定成败的不是会不会写代码,而是懂不懂业务。营销团队也该让 Agent 接走重复性工作,把人还给创意与策略。 2026-07-07 Reading the Rhythm of AI Work: What the Anthropic Economic Index Cadences Report Reveals ◐ ECON-1 The Anthropic Economic Index Cadences report samples Claude usage hourly, exposing when people turn to AI and what they produce. Marketers can read the rhythm to time content, staff support, and budget compute where value density is highest. 2026-07-07 Why Model Specialization Is Inevitable ◐ NICHE-1 No free lunch has held for thirty years: generalists do not win, specialists do. Choosing a model is a trade-off between compute, brand voice, and data compliance—here is a framework. 2026-07-07 用 RAG 给营销人装一个不会胡说的脑子 ◐ RAG-7 大模型会胡说,但营销文案不能。RAG(检索增强生成)是把品牌知识库塞进模型、让它只说真话的关键一步——本文讲营销人为什么需要它、怎么落地。 2026-07-07 Why Brand Voice Holds or Slips Inside Claude ◐ COGNITION-1 The same brand brief can yield a coherent campaign or a disjointed string of sentences. Anthropic has found a small cluster of neurons inside Claude that acts as a central broadcast station, explaining the gap and showing how to make brand context actually stick. 2026-07-07 When Agents Work for Hours, Not Seconds: Inside OpenAI's Findings on Transformed Work ◐ AGENT-1 OpenAI internal data shows agents now run tasks spanning minutes to hours, with a quarter of Codex requests corresponding to over eight hours of human work. The findings mark a shift from chat to delegated long-horizon work. 2026-07-07 The RAG Bill — Where the Money Goes in a Retrieval Pipeline ◐ THRIFT-1 Last stop showed how RAG lets a brand knowledge base answer open-book. This time we land on THRIFT-1 to dissect where the money goes in a single RAG query — embedding, vector store, LLM inference — and the hardware and quantization levers that actually cut cost. 2026-07-07 Autonomous Research at Scale: Inside OpenAI's Deep Research ◐ RESEARCH-1 OpenAI's Deep Research runs a single prompt through 5 to 30 minutes of multi-step browsing, backtracking, and synthesis to produce a cited report. This log unpacks how the agent actually works, where it breaks, and what it means for marketing insight teams. 2026-07-07 一句话长出一个品牌知识站:AI 内容工厂的新玩法 ◐ SITE-1 AI agent skill 让你一句话生成一个带 SEO、PWA、闪卡、测验的完整知识网站——营销人该看的玩法与红线。 2026-07-07 给 AI 装上"浏览器":营销人的 24 小时全网监测员 ◐ ACCESS-1 一个让 AI 直接访问网页的 agent skill——读竞品页、盯价格、抓评论、追热点,都让 agent 替你跑腿。但红线也得先画清楚。 2026-07-07 Embeddings, for marketers who skipped the math ◐ VECTOR-1 Every "AI for marketers" pitch eventually says "embedding." This log entry explains what an embedding actually is — without calculus — and why it's the engine under semantic search, recommendations, and RAG.