2026-09-22 AI 日报
统计日时区:America/Los_Angeles · 已结算
google/ax · dream-num/univer · Asymptote-Labs/agent-beacon
开源项目
- google/ax
Google's open agentic orchestration runtime
- dream-num/univer
The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
- Asymptote-Labs/agent-beacon
The cross-harness self-improving memory layer for AI agents.
- browserbase/stagehand
The SDK to extract data and interact with any site on the web. Get started with Claude Code, Codex, Eve, Mastra, and more.
- can1357/oh-my-pi
⌥ Coding agent with the IDE wired in
- browser-use/video-use
Edit videos with coding agents
- stablyai/orca
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
- davila7/claude-code-templates
CLI tool for configuring and monitoring Claude Code
- obra/superpowers
An agentic skills framework & software development methodology that works.
- DeusData/codebase-memory-mcp
High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms…
论文
- WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across…
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the…
- Realtime-Venus: A full-duplex interaction system with asynchronous delegation
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and…
- GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and…
- OmniEdu: Open Foundation Models for Learning and Teaching
Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional…
- Grounded Action Model: 3D Grounding as a Foundation for Robotics
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from…
- Transferring the Intelligence of VLMs to Robotic Control
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human…
- onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its…
- One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents
Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven:…
- Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents,…
行业新闻
- GPT-6 Sol and Luna
- Claude Opus 5.5
We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less…
- OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
- OpenAI is well positioned to fast-follow Jev
- Claude Opus 5.5
https://github.com/anthropics/ClaudeForFoundationModels/comm...
- Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
- LLM Ass Bench
- People Training OpenAI's AI Fired for Using AI to Train the AI
- Show HN: JevBench, a reproducible benchmark for typed decision models
Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in…
- The new CC, an AI agent built for families
Google Labs introduces a new experimental agent built for families, helping households spend less time on logistics and more time together. Google is…
社区动态
- Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths…
- Claude Opus 5.5 is available today.
Claude Opus 5.5 is available today.
- GPT-6 Sol and Luna just landed in Astra’s orbit.
GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and…
- Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to…
- We want the OpenAI API to feature the best model at every price point and to be the best at every modality (text,…
We want the OpenAI API to feature the best model at every price point and to be the best at every modality (text, code, image, video, etc). And then we want…
- Grok @Bot now in your Tesla!
Grok @Bot now in your Tesla!
- As part of our efforts to pace the frontier, we’re committed to supporting independent assessments with deep levels of…
As part of our efforts to pace the frontier, we’re committed to supporting independent assessments with deep levels of access across training, evaluation, and…
- Introducing GPT-6 Sol and Luna
- Introducing Claude Opus 5.5, 40% Cheaper and Smarter Than Ever Before
- Pirate Face - pirate bay for LLMs
The title says for itself In case someone desides to censor huggingface, we'll have an alternative
官方发布
- Introducing GPT-6 Sol and Luna
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
- Introducing Grok 4.7
- Better prompt caching for GPT-6
Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
- On Claude Opus 5.5, thinking can't be disabled: thinking: {"type": "disabled"} and thinking: {"type": "enabled", ...}…
On Claude Opus 5.5, thinking can't be disabled: thinking: {"type": "disabled"} and thinking: {"type": "enabled", ...} return a 400 error. Omit the thinking…
- Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation
We are excited to open-source Qwen-Image-2.1, an image model in the Qwen family that balances generation quality, inference efficiency, and cost.…
- ChatGPT Ads expands to Southeast Asia and Taiwan
ChatGPT Ads is expanding to Southeast Asia and Taiwan, giving eligible businesses new ways to reach people across more than 60 countries.
- Introducing Grok Voice Transcribe 2.0
- Qwen/Qwen-Image-2.1-PE-T2I
text-to-image
- Tools can now be defined inside a mid-conversation system message, in beta on the Claude API with the…
Tools can now be defined inside a mid-conversation system message, in beta on the Claude API with the inline-tools-2026-09-15 beta header. A tool_addition…
- Qwen3.8-LiveTranslate: Names the speaker. Carries the meaning.
Simultaneous interpretation is not only about translating fast — it must also hear clearly and translate accurately. Qwen3.8-LiveTranslate rebuilds real-time…