2026-09-24 AI 日报
统计日时区:America/Los_Angeles · 尚未结算
androoAGI/starnet · Asymptote-Labs/agent-beacon · NVIDIA/Model-Optimizer
开源项目
- androoAGI/starnet
A living pixel-art station where real AI agents do real work. Local-first desktop agent harness - bring your own key, watch your crew actually run.
- Asymptote-Labs/agent-beacon
The cross-harness self-improving memory layer for AI agents.
- NVIDIA/Model-Optimizer
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It…
- paperclipai/paperclip
The open-source app everyone uses to manage agents at work
- google/ax
Google's open agentic orchestration runtime
- rohitg00/ai-engineering-from-scratch
Learn it. Build it. Ship it for others.
- mvschwarz/openrig
Multi-agent harness that runs Claude Code and Codex together as one system
- dream-num/univer
The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
- stablyai/orca
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
- strands-agents/harness-sdk
Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
论文
- SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who…
- Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs…
- The Past Frames the Future: Memory for Autoregressive Video Generation
Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments.…
- HappyWorld-Bench
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration,…
- RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling
Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing…
- Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed,…
- Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built…
- PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing
Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing…
- MemBodied: Recurrent Associative Memory for Vision-Language-Action Models
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage…
- PACT: From Credit Assignment to Critic Alignment
Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted…
行业新闻
- Best LLM for every budget, updated daily
- Hackers influence ChatGPT and Gemini to direct users to scam centers
- Using LLMs to trace alchemical knowledge and decode 17th century letters
- MentalHealthBench
- Humans Are Reading Your ChatGPT Chats, Lawsuit Claims
- GPT-6 Sol is like GPT-5.6 Terra, GPT-6 Luna is like GPT-5.6 Asteroid
- Docker releases cloud sandboxes, enabling safe agentic workloads in the cloud
- A Million Agents Is a Distributed System Problem
- Show HN: Radix – Visual UI for agentic programming
Hey HN, I'm Jordan from Radix. Radix is a UI tool for programming agents. You prompt your agent to generate a workspace for a task you're working on and get…
- Paul Graham on LLMs 'Thinking'
社区动态
- @vasalex93 1.
@vasalex93 1. We will keep accelerating. Our AI efforts are only 3 years old, vs 6 and 10 years old for Anthropic and OpenAI. If our second derivative remains…
- Situation today: GPT-6 Sol is like GPT-5.6 Terra GPT-6 Luna is like GPT-5.6 Asteroid https://t.co/D7OiG6KI8Q
Situation today: GPT-6 Sol is like GPT-5.6 Terra GPT-6 Luna is like GPT-5.6 Asteroid https://t.co/D7OiG6KI8Q
- I benchmarked repowise, CodeGraph, Serena, Graphify, code-review-graph, cocoindex and codebase-memory-mcp across Codex, Claude Code and a local model. The 60-90% token-saving claims didn't hold up
Scroll to bottom for tldr In July, JetBrains reran the headline claims of two token-saving tools on real agent workloads. Caveman claimed 65% and measured…
- Publication venue recommendations [R]
I am 5th year and have no published research so far. All my papers have been consistently rejected from top tier AI conferences even though they got good…
- Qwen-3.8-27B is good enough that I stopped using API
Many a praise have been sung on Qwen-3.8, but here is mine. Qwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painful, so…
- UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy
Hey everyone, Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens…
- Qwen 3.8 27b be like...
The user is frustrated — I rambled too much and didn't act. Let's just run the test suite and move on. No more forensics. One command, execute, then report.…
- What's up with AAAI reviewers and organizers? [D]
My paper advanced to the second round...but... Out of the papers I reviewed. One did not follow the AAAI template and was unblinded. My review was two lines.…
- Jive - Rethinking the Agentic Loop with System One Models
I have been thinking that the current Agentic Loop design of LLM Call -> Tool Call -> ... has been outdated. The arrival of Jev and other System One models…
- We have a new acronym for you: SAFA.
We have a new acronym for you: SAFA. The Google, OpenAI and Anthropic AI safety body is starting to take shape, with a tentative name of the Standards…
官方发布
- Coding sessions are longer and use more context. Claude Opus 5.5 is built with that in mind.
- Introducing Grok 4.7
- Introducing GPT-6 Sol and Luna
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
- Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation
We are excited to open-source Qwen-Image-2.1, an image model in the Qwen family that balances generation quality, inference efficiency, and cost.…
- Introducing Grok Voice Transcribe 2.0
- Gemini 3.8 text-to-speech says hello
Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.
- Turn your REST APIs into MCP tools with Google Cloud API Gateway
Google Cloud API Gateway now acts as a native remote Model Context Protocol (MCP) server, eliminating the need to build and maintain custom middleware to…
- Claude discovers a novel enzyme system with CRISPR-like repeats
We’re announcing a new life sciences research group and laboratory at Anthropic. This post introduces the team behind this work and shares early results in…
- Better prompt caching for GPT-6
Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
- Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs
The MaxText team successfully reproduced AI2’s OLMo 3 7B language model from scratch on Google Cloud TPUs using JAX/XLA, precisely matching the original…