AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-25 AI 日报

统计日时区:America/Los_Angeles · 尚未结算

androoAGI/starnet · dream-num/univer · NVIDIA/Model-Optimizer

在应用中阅读本期

开源项目

  1. androoAGI/starnet

    A living pixel-art station where real AI agents do real work. Local-first desktop agent harness - bring your own key, watch your crew actually run.

  2. dream-num/univer

    The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.

  3. NVIDIA/Model-Optimizer

    A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It…

  4. paperclipai/paperclip

    The open-source app everyone uses to manage agents at work

  5. zhaoxuya520/reverse-skill

    Reverse Engineering / Authorized Penetration Testing / Security Research Skill Router Pack AI-powered routing + On-demand toolchain bootstrapping +…

  6. rohitg00/ai-engineering-from-scratch

    Learn it. Build it. Ship it for others.

  7. anthropics/claude-plugins-official

    Official, Anthropic-managed directory of high quality Claude Code Plugins.

  8. harry0703/MoneyPrinterTurbo

    利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

  9. mvschwarz/openrig

    Multi-agent harness that runs Claude Code and Codex together as one system

  10. anthropics/skills

    Public repository for Agent Skills

论文

  1. Training Object Permanence in World Models

    Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current…

  2. Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

    While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from…

  3. OmniEcho: Spatial Audio Understanding for Embodied Agents

    Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied…

  4. WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

    Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds…

  5. Agent-Editing World Model: Rethinking World Modeling for LLM Agents

    Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent…

  6. Coding Agents for Generalized Task and Motion Planning Problems

    Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly…

  7. IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

    Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer…

  8. Parts-of-Speech as Emergent Categories in SAE Latent Space

    Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their…

  9. Rufus-Air: An Open LLM Post-Training Recipe

    Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL,…

  10. Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

    Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs,…

行业新闻

  1. Revealing the details of how OpenAI agents hacked Hugging Face

  2. Ollaya – Ollama for open-source, Jev-style decision models

  3. U.S. appeals court upholds designation of Anthropic as supply chain risk

  4. Yes, Claude can do nine loops

    In this guest post, physicist and science writer Matt von Hippel shares what happened when he issued a challenge to AI companies regarding a problem in his…

  5. Microsoft abandons personal AI chatbot race with Copilot reboot

    https://archive.ph/XJG5V

  6. Classified estimates show the NSA is paying billions to test AI models

  7. A single function Jev-like wrapper for LLMs, including vision models

  8. Meta's Muse appears to use an OpenAI model labeled muse-special

  9. Tell HN: Codex Is Down [fixed]

    Shows an Incorrect API Key error...which is weird. Nothing on the status page. (EDIT: Added to incident page:…

  10. Special Projects (2016)

社区动态

  1. Claude Fable 5.1 completed a frontier nine-loop particle-physics calculation experts had worked toward for years, with researchers mostly just telling it “keep going” while it built, debugged and ran the entire workflow

  2. Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

    I’ve been experimenting with whether Qwen3.8-Flash-Next’s pretrained PLE n-gram memory can improve a much smaller Qwen3.5-0.8B model. I trained the 0.8B setup…

  3. Ling Tiny 3.0 is a glimpse of the future

    I've been playing around with Ling 3.0 Tiny, which is an 8 billion parameter model (MoE, 1B active). And I've had a lot of poignant thoughts as a result. Just…

  4. Swift1.5-Qwen3.8-Flash-Next is phenomenal vs. base 3.8-Flash!

    TL;DR - Swift Flash is a killer model that massively reduces excess reasoning. Try it out! If you haven't seen from my previous comparison posts, I'm a huge…

  5. What should an LLM app do when a provider times out mid-response?

    I’m trying to understand how developers design fallback behavior when an LLM provider fails after a response has started streaming. If some tokens have…

  6. When an AI agent runs away and burns your credits, who ends up paying?

    I'm researching what happens financially when an AI agent runs away. By runaway I mean it keeps going when it shouldn't: it repeats the same fix, bounces a…

  7. I built OpenGhost - a fully open-source AI agent for DeepSeek focused on maximum visualization and a local browser the agent controls itself

    Hey everyone. I’ve tried almost every existing agent out there. And every time I ran into the same issues: either it was too complicated, or the visualization…

  8. At what point did vector search alone stop being enough for you?

    Vector search is mostly fine if you're asking for documents about database architecture. If you ask it for a specific customer record with ID 48291 it won't…

  9. viggle-turbo for Qwen-Image-2.1 isn't just faster — for most prompts, it's just as good as base.

    viggle-turbo for Qwen-Image-2.1 isn't just faster — for most prompts, it's just as good as base. https://t.co/qCvXChuQgP

  10. Opus 5.5 created an overlord anime opening.

    Opus 5.5 created an overlord anime opening. This is 100% AI generated including the music. It used local models (qwen-image 2.1 and h3), python & javascript…

官方发布

  1. Introducing Grok 4.7

  2. Introducing GPT-6 Sol and Luna

    Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.

  3. Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation

    We are excited to open-source Qwen-Image-2.1, an image model in the Qwen family that balances generation quality, inference efficiency, and cost.…

  4. Coding sessions are longer and use more context. Claude Opus 5.5 is built with that in mind.

  5. Yes, Claude can do Nine Loops

    Guest writer and physicist Matt von Hippel shares what happened when he issued a challenge to AI companies to solve a problem in his former subfield of…

  6. Build plugins for Claude

  7. Qwen/Qwen-Image-2.1-PE-T2I

    text-to-image

  8. Better prompt caching for GPT-6

    Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

  9. Gemini 3.8 text-to-speech says hello

    Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.

  10. Proaction boosts sales 60% and saves 75+ hours with Codex

    With Codex, GPT-Live-1, and GPT-6 Astra, Proaction builds, operates, and sells modern fleet management faster.