AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-24 AI 日报

统计日时区:America/Los_Angeles · 尚未结算

androoAGI/starnet · Asymptote-Labs/agent-beacon · NVIDIA/Model-Optimizer

在应用中阅读本期

开源项目

  1. androoAGI/starnet

    A living pixel-art station where real AI agents do real work. Local-first desktop agent harness - bring your own key, watch your crew actually run.

  2. Asymptote-Labs/agent-beacon

    The cross-harness self-improving memory layer for AI agents.

  3. NVIDIA/Model-Optimizer

    A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It…

  4. paperclipai/paperclip

    The open-source app everyone uses to manage agents at work

  5. google/ax

    Google's open agentic orchestration runtime

  6. rohitg00/ai-engineering-from-scratch

    Learn it. Build it. Ship it for others.

  7. mvschwarz/openrig

    Multi-agent harness that runs Claude Code and Codex together as one system

  8. dream-num/univer

    The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.

  9. stablyai/orca

    Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.

  10. strands-agents/harness-sdk

    Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.

论文

  1. SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

    Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who…

  2. Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

    Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs…

  3. The Past Frames the Future: Memory for Autoregressive Video Generation

    Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments.…

  4. HappyWorld-Bench

    Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration,…

  5. RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

    Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing…

  6. Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

    Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed,…

  7. Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?

    Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built…

  8. PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

    Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing…

  9. MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

    Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage…

  10. PACT: From Credit Assignment to Critic Alignment

    Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted…

行业新闻

  1. Best LLM for every budget, updated daily

  2. Hackers influence ChatGPT and Gemini to direct users to scam centers

  3. Using LLMs to trace alchemical knowledge and decode 17th century letters

  4. MentalHealthBench

  5. Humans Are Reading Your ChatGPT Chats, Lawsuit Claims

  6. GPT-6 Sol is like GPT-5.6 Terra, GPT-6 Luna is like GPT-5.6 Asteroid

  7. Docker releases cloud sandboxes, enabling safe agentic workloads in the cloud

  8. A Million Agents Is a Distributed System Problem

  9. Show HN: Radix – Visual UI for agentic programming

    Hey HN, I'm Jordan from Radix. Radix is a UI tool for programming agents. You prompt your agent to generate a workspace for a task you're working on and get…

  10. Paul Graham on LLMs 'Thinking'

社区动态

  1. @vasalex93 1.

    @vasalex93 1. We will keep accelerating. Our AI efforts are only 3 years old, vs 6 and 10 years old for Anthropic and OpenAI. If our second derivative remains…

  2. Situation today: GPT-6 Sol is like GPT-5.6 Terra GPT-6 Luna is like GPT-5.6 Asteroid https://t.co/D7OiG6KI8Q

    Situation today: GPT-6 Sol is like GPT-5.6 Terra GPT-6 Luna is like GPT-5.6 Asteroid https://t.co/D7OiG6KI8Q

  3. I benchmarked repowise, CodeGraph, Serena, Graphify, code-review-graph, cocoindex and codebase-memory-mcp across Codex, Claude Code and a local model. The 60-90% token-saving claims didn't hold up

    Scroll to bottom for tldr In July, JetBrains reran the headline claims of two token-saving tools on real agent workloads. Caveman claimed 65% and measured…

  4. Publication venue recommendations [R]

    I am 5th year and have no published research so far. All my papers have been consistently rejected from top tier AI conferences even though they got good…

  5. Qwen-3.8-27B is good enough that I stopped using API

    Many a praise have been sung on Qwen-3.8, but here is mine. Qwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painful, so…

  6. UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy

    Hey everyone, Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens…

  7. Qwen 3.8 27b be like...

    The user is frustrated — I rambled too much and didn't act. Let's just run the test suite and move on. No more forensics. One command, execute, then report.…

  8. What's up with AAAI reviewers and organizers? [D]

    My paper advanced to the second round...but... Out of the papers I reviewed. One did not follow the AAAI template and was unblinded. My review was two lines.…

  9. Jive - Rethinking the Agentic Loop with System One Models

    I have been thinking that the current Agentic Loop design of LLM Call -> Tool Call -> ... has been outdated. The arrival of Jev and other System One models…

  10. We have a new acronym for you: SAFA.

    We have a new acronym for you: SAFA. The Google, OpenAI and Anthropic AI safety body is starting to take shape, with a tentative name of the Standards…

官方发布

  1. Coding sessions are longer and use more context. Claude Opus 5.5 is built with that in mind.

  2. Introducing Grok 4.7

  3. Introducing GPT-6 Sol and Luna

    Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.

  4. Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation

    We are excited to open-source Qwen-Image-2.1, an image model in the Qwen family that balances generation quality, inference efficiency, and cost.…

  5. Introducing Grok Voice Transcribe 2.0

  6. Gemini 3.8 text-to-speech says hello

    Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.

  7. Turn your REST APIs into MCP tools with Google Cloud API Gateway

    Google Cloud API Gateway now acts as a native remote Model Context Protocol (MCP) server, eliminating the need to build and maintain custom middleware to…

  8. Claude discovers a novel enzyme system with CRISPR-like repeats

    We’re announcing a new life sciences research group and laboratory at Anthropic. This post introduces the team behind this work and shares early results in…

  9. Better prompt caching for GPT-6

    Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

  10. Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

    The MaxText team successfully reproduced AI2’s OLMo 3 7B language model from scratch on Google Cloud TPUs using JAX/XLA, precisely matching the original…