2026-09-30 AI 日报
统计日时区:America/Los_Angeles · 已结算
Niko1221/Strata · magnitudedev/magnitude · NVIDIA/OpenShell
开源项目
- Niko1221/Strata
Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image…
- magnitudedev/magnitude
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to…
- NVIDIA/OpenShell
OpenShell is the safe, private runtime for autonomous AI agents.
- ifixai-ai/iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is…
- mvschwarz/openrig
Build your own network of agents from Claude Code, Codex and Pi: persistent teams with roles, shared context and owned work.
- DietrichGebert/ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- mksglu/context-mode
Context window optimization for AI coding agents. Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via…
- t8y2/dbx
25 MB lightweight cross-platform database client for 100+ databases, including MySQL, PostgreSQL, SQLite, Redis, MongoDB, DuckDB, SQL Server, and Dameng.…
- diegosouzapw/OmniRoute
Never stop coding. Free MIT AI gateway: one endpoint, 359 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with…
- heygen-com/hyperframes
Write HTML. Render video. Built for agents.
论文
- Raven: The Harness of Harnesses for Composable Agentic Intelligence
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition…
- Omni-IO Skills: Harnessing Your Agent Omni-Native
General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video,…
- What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must…
- In-Context Learning for Robots: Methods and Applications
General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for…
- MaLiang-Harness: A Programmable Path to Image and Video Generation
Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation.…
- PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual…
- LongLive-Plug: Once-for-All Distillation for Video Generation
Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation…
- Think Before You Score: Thinking Reward Model for Visual Generation
Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate…
- Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and…
- Scaling Properties of Same-Family On-Policy Distillation
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across…
行业新闻
- Gemini 4 Argon
See also: Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
- Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It…
- You said no MCP
- Gemini 4 Argon (High): Intelligence, Performance and Price Analysis
See also: Gemini 4 Argon - https://news.ycombinator.com/item?id=49913571
- Doing a Machine Learning PhD While Working in Japan
- GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
- Claude Says
- Anthropic's IPO Prospectus Is a Fucking Doozy
- Is sandboxing sufficient to contain rogue agents?
- FTC opens probe into AI giants including Anthropic and OpenAI
社区动态
- Introducing Gemini 4 Argon – our new frontier model.
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense –…
- Announcing Gemini 4 Argon, our new frontier model.
Announcing Gemini 4 Argon, our new frontier model. Argon is built to sustain deep reasoning across complex, long-horizon workflows and delivers frontier…
- Introducing Gemini 4 Argon, our new frontier model, rolling out to cyber defenders starting today, and more widely as…
Introducing Gemini 4 Argon, our new frontier model, rolling out to cyber defenders starting today, and more widely as soon as possible. I am really excited by…
- Gemini 4 Argon has an insanely low hallucination rate on Artificial Analysis.
Gemini 4 Argon has an insanely low hallucination rate on Artificial Analysis. 15%. Grok 4.7 is at 29%. GPT-6 Astra 45%. Opus 5.5 59%. Fable 5.1 69%. The only…
- Gemini 4 Argon: our next era of frontier intelligence
Cyber - so locked down to select partners. Decent benchmarks.
- Open source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU.
- add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp
now you can use GLM-5.3-Flash on your home computer
- Built on MiniMax H3, @Creatify_Labs' Boreal-H3 is a video model optimized for advertising, keeping products and…
Built on MiniMax H3, @Creatify_Labs' Boreal-H3 is a video model optimized for advertising, keeping products and characters consistent while following creative…
- Small teams are taking on more with AI—from finding customers to building products and managing finances.
Small teams are taking on more with AI—from finding customers to building products and managing finances. Our new report explores how small businesses are…
- Impressive work by the @HeyGen on the launch of HeyGen Video!
Impressive work by the @HeyGen on the launch of HeyGen Video! ✨ Built on MiniMax H3 and post-trained by HeyGen, it brings production-quality video to…
官方发布
- Gemini 4 Argon: our next era of frontier intelligence
Announcing Gemini 4 Argon, our frontier model for real-world coding, enterprise knowledge work, and cyber defense, rolling out soon.
- Introducing GPT-6.1 Sol
Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.
- Introducing dots
Dots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work…
- Claude for Government is now generally available
- Basis completes a tax workbook 2x faster with GPT-6 Astra
GPT-6 Astra completed a 50-tab tax workbook twice as fast as GPT-5.6 Sol, and its stronger understanding of user intent gives Basis more confidence in…
- Let skills in Gemini tackle your most repetitive tasks
Now you can automate your most repetitive tasks more easily with reusable custom instructions through skills, which will be replacing gems.
- Coding sessions are longer and use more context. Claude Opus 5.5 is built with that in mind.
- Z.ai GLM 5.3 (zai-glm-5-3) is now Generally Available.
Z.ai GLM 5.3 (zai-glm-5-3) is now Generally Available.
- What work can robots do?
We built an index of how well today’s robots can perform US job tasks. Robots can already do three-quarters of physical tasks, mostly in limited settings, but…
- Introducing SynthID Bio
Proof of concept for watermarking AI-generated proteins while preserving biological function.