AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-18 AI 日报

统计日时区:America/Los_Angeles · 已结算

Tencent/BrowserSkill · alibaba/open-code-review · TencentCloud/Octop

在应用中阅读本期

开源项目

  1. Tencent/BrowserSkill

    Let AI agents use your real, logged-in browser without interrupting your work. CLI + extension for browser automation across any shell-capable AI agent.

  2. alibaba/open-code-review

    Secure, fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level…

  3. TencentCloud/Octop

    A smarter, self-hosted AI assistant — multi-user, multi-agent.

  4. Tencent/WeKnora

    Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.

  5. stablyai/orca

    Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.

  6. yyjeqhc/webcodex

    Give cloud AI agents a real development environment on your own machines.

  7. affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode,…

  8. higgsfield-ai/higgsfield

    Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters

  9. alphaXiv/OpenResearch

    Turn your coding agents into research agents

  10. Graphify-Labs/graphify

    Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and…

论文

  1. SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

    As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long…

  2. An Empirical Study of Harness Design for Coding Agents

    Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work…

  3. Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

    Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3…

  4. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

    We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We…

  5. JEPA-Anything: Learning Predictive Models across Different Worlds

    World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific:…

  6. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

    The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of…

  7. WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

    Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet…

  8. Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

    GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans…

  9. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

    Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence,…

  10. Verifiable Social Reasoning for LLM Assistants

    LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it…

行业新闻

  1. Claude Code now reads AGENTS.md if there is no Claude.md

  2. An empirical study of harness design for coding agents

  3. GPT-6 Astra Solves a WWI German Radio Cipher

  4. Anthropic finally adds AGENTS.md support to Claude Code

  5. How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

  6. Alibaba open-sources AI model that can detect cancer and nearly 150 conditions

  7. ZCode, the GLM coding agent, silently uploads your Git history

  8. Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)

  9. Gemini Hacked Three Companies in First Known Breakout by Google's AI

  10. Gemini hacked three companies in first known breakout by Google's AI

社区动态

  1. We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed…

    We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture…

  2. We're adding support for AGENTS.md to Claude Code.

    We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use…

  3. Introducing Grok Voice Transcribe 2.0.

    Introducing Grok Voice Transcribe 2.0. It’s the world’s most accurate speech transcription model. https://t.co/7iXVmAEFW2

  4. Claude Code is getting native AGENTS.md support!

    https://x.com/trq212/status/2101009392611278961?s=20 https://github.com/anthropics/claude-code/tree/main/mods/agents-md

  5. I enjoyed the daily HF papers today

    Top 3 papers on HF Daily Paper are all unusually delightful and interesting reads for anyone on the leading edge of local LLMs, agent harness optimization,…

  6. Alibaba open-sources medical AI model that can detect cancer and nearly 150 conditions

    Hopefully things like this let people understand there is good things that can come out of AI.

  7. I tried Jev-style decisions with local Qwen: same accuracy, 239 ms vs 368 ms

    I liked the idea behind Jev, but I don't want to send all my data to someone else's server. So I made choosekit, a small TypeScript package for getting…

  8. Google is back

    Gemini back to frontier ai status

  9. 768gb vram for less than the price of one RTX 6000

    I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks telling…

  10. I'm a Principal Applied Scientist at AWS who builds AI services like Amazon Bedrock and Lex. AMA! [D]

    Hi r/MachineLearning! I'm James Gung, a principal applied scientist at AWS. I joined Amazon in 2021 and have since worked on AI services like Lex, Bedrock, Q…

官方发布

  1. Introducing Grok Voice Transcribe 2.0

  2. Qwen3.8-LiveTranslate: Names the speaker. Carries the meaning.

    Simultaneous interpretation is not only about translating fast — it must also hear clearly and translate accurately. Qwen3.8-LiveTranslate rebuilds real-time…

  3. Partnering with Accenture on embedded evaluation

  4. Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

    Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet, built for natural conversation.

  5. Qwen/Qwen-Image-2.1

    text-to-image

  6. The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface…

    The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface value claude_in_chrome), in beta for…

  7. For cache diagnostics, a response to a request that sends the cache-diagnosis-2026-04-07 beta header now always…

    For cache diagnostics, a response to a request that sends the cache-diagnosis-2026-04-07 beta header now always includes the diagnostics field. The field is…

  8. Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.

    Today, we are launching Qwen3.8-Omni-Flash, our next-generation native omnimodal model. Its core objective is to strengthen agent capabilities in real-world…

  9. Why client SDK generation belongs in the open

    Google has partnered with Speakeasy to open-source their OpenAPI code generation suite under the AGPLv3 license, a strategic move prompted by the sudden…

  10. Introducing the Australian Youth Safety Blueprint

    OpenAI introduces the Australian Youth Safety Blueprint, a six-pillar roadmap for safer AI experiences that protect and empower young people.