2026-10-01 AI 日报
统计日时区:America/Los_Angeles · 已结算
mvschwarz/openrig · Panniantong/Agent-Reach · ifixai-ai/iFixAi
开源项目
- mvschwarz/openrig
Build your own network of agents from Claude Code, Codex and Pi: persistent teams with roles, shared context and owned work.
- Panniantong/Agent-Reach
Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
- ifixai-ai/iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is…
- DietrichGebert/ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- NVIDIA/OpenShell
OpenShell is the safe, private runtime for autonomous AI agents.
- earendil-works/pi
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
- heygen-com/hyperframes
Write HTML. Render video. Built for agents.
- magnitudedev/magnitude
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to…
- JuliusBrussee/caveman
🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
- FalkorDB/FalkorDB
A super fast Graph Database uses GraphBLAS under the hood for its sparse adjacency matrix graph representation. Our goal is to provide the best Knowledge…
论文
- Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than…
- AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test…
- The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial…
- RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable…
- Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies
GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities…
- Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle…
- UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback.…
- WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in…
- False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This…
- EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps,…
行业新闻
- Clef: Open-weight decision models, and new RL fine-tuning platform
- DeepSeek Harness Desktop for macOS and Windows
Install plugins, or create them through chat in “Creator mode”, to extend tools, skills, and the interface. The floating Pomodoro timer is built, installed,…
- RIP, vector database
- FTC is investigating OpenAI, Anthropic and other AI companies over product risks
- ArXiv's Updated Rate Limit Policy
- Context Language Models
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that…
- Figma restricts MCP access to whitelisted clients, excluding Pi
- GPT-Synopsys: Frontier Intelligence to Revolutionize Chip Design
- Identity Management for Agentic AI [pdf] (2025)
- Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia
社区动态
- You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a…
You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude…
- In physics, an “impedance mismatch” occurs when two systems each work well but are poorly matched.
In physics, an “impedance mismatch” occurs when two systems each work well but are poorly matched. In this Science Blog guest post, Harvard physicist Matthew…
- SynthID Bio is our new family of watermarking methods made for AI-generated biological designs.
SynthID Bio is our new family of watermarking methods made for AI-generated biological designs. In a world first, we can now embed an imperceptible signature…
- arXiv now limits submitters to up to two submissions per calendar month [N]
- Grok 4.7 is now available on the Gemini Enterprise Agent Platform https://t.co/AdWgDF8Y2v
Grok 4.7 is now available on the Gemini Enterprise Agent Platform https://t.co/AdWgDF8Y2v
- Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp
now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B? (merged after 17h of development) quants:…
- Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now
pi-llama-skip-reasoning is an extension for the Pi.dev harness that forces a local llama.cpp model to stop reasoning and answer / act immediately. When you…
- The AI industry has discovered intellectual property
OpenAI says Moonshot-linked operators used thousands of accounts to extract protected reasoning from its models for adversarial distillation. No encryption…
- Now in Grok Build: A new Agent Dashboard.
Now in Grok Build: A new Agent Dashboard. All your agents on one screen. Try it with /dashboard https://t.co/ycsvWoYT0I
- A Harvard physicist spent 3 months doing research with Claude Fable 5: it reproduced weeks of work in 20 minutes, completed 15 never-before-solved physics calculations, and contributed to 36 papers across 18 fields
官方发布
- Gemini 4 Argon: our next era of frontier intelligence
Announcing Gemini 4 Argon, our frontier model for real-world coding, enterprise knowledge work, and cyber defense, rolling out soon.
- Introducing GPT-6.1 Sol
Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.
- Introducing dots
Dots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work…
- We've launched Claude Sonnet 5.5 (claude-sonnet-5-5).
We've launched Claude Sonnet 5.5 (claude-sonnet-5-5). It's available on the Claude API, Claude in Amazon Bedrock, Claude Platform on AWS, Claude on Google…
- Basis completes a tax workbook 2x faster with GPT-6 Astra
GPT-6 Astra completed a 50-tab tax workbook twice as fast as GPT-5.6 Sol, and its stronger understanding of user intent gives Basis more confidence in…
- Claude-shaped science
Guest author Prof. Matthew Schwartz describes what happened when he stopped fighting Claude and allowed Claude to find “Claude-shaped” problems: ones best…
- Z.ai GLM 5.3 (zai-glm-5-3) is now Generally Available.
Z.ai GLM 5.3 (zai-glm-5-3) is now Generally Available.
- Customize Claude Code with mods
- Guided Vision in Gemini Live: built for accessibility
Guided Vision in Gemini Live is built alongside the blind and low-vision community and offers real-time visual assistance.
- Let skills in Gemini tackle your most repetitive tasks
Now you can automate your most repetitive tasks more easily with reusable custom instructions through skills, which will be replacing gems.