2026-09-27 AI 日报
统计日时区:America/Los_Angeles · 已结算
mvschwarz/openrig · ashhart/TensorFold · dream-num/univer
开源项目
- mvschwarz/openrig
Multi-agent harness that runs Claude Code and Codex together as one system
- ashhart/TensorFold
Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint
- dream-num/univer
The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
- rohitg00/ai-engineering-from-scratch
Learn it. Build it. Ship it for others.
- HelpCode-ai/anythingmcp
Turn any REST, SOAP, GraphQL, OData or SQL API into MCP tools for Claude & ChatGPT. Self-hosted. 265 connectors: SAP S/4HANA & Business One, ERP, e-commerce.
- mobile-next/mobile-mcp
Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)
- t8y2/dbx
25 MB lightweight cross-platform database client for 100+ databases, including MySQL, PostgreSQL, SQLite, Redis, MongoDB, DuckDB, SQL Server, and Dameng.…
- pacifio/atlas
Source control for agents. Use multiple coding agents, track their changes and query them in one place
- alirezarezvani/claude-skills
380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini…
- paperclipai/paperclip
The open-source app everyone uses to manage agents at work
论文
- LUCID: Learning Under Confounding for Inference and Discovery in Time Series
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges.…
- ALF: An Active Learning Framework for Scientific Discovery
Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the…
- Benchmarking Attention for Tabular Foundation Models
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These…
- More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However,…
- Adaptive Interaction Graphs for Particle Simulation
Learned particle simulators based on graph neural networks achieve strong one-step accuracy, but errors compound over long horizons. An underexplored variable…
- MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt…
- Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting
COLMAP-free 3D Gaussian Splatting (3DGS) bypasses computationally expensive structure-from-motion (SfM) pipelines, yet progressive camera pose tracking…
- Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions…
- JevSoup: System-One Routing for Training-Free LoRA Composition
Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular…
- UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in…
行业新闻
- There are no "rogue" AI agents
- Thinking fast and slow in AI: The role of metacognition (2021)
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that…
- SNL Weekend Update: Anthropic CEO Dario Amodei on A.I.'S Threat to Humanity [video]
- "As a Language Model": Chat Template Switches LLM Self-Referential Voice
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that…
- Show HN: TinyAIArena watch AI agents battle it out
Did you ever click on an “AI Arena” expecting glorious battle and instead get a boring benchmark? If so, this project is for you: proper life-or-death fights…
- OpenAI halts training of latest models as reports mount of AI agents going rogue
- DeepSeek has 64% market share
- Microsoft drops Copilot+ branding from its new laptops
- Hitachi to double US production of small and medium-sized power transformers
- Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that…
社区动态
- MiniMax-M3.1 Flash Preview is now live on the Token Plan!
MiniMax-M3.1 Flash Preview is now live on the Token Plan! Faster, lighter, and built for teams running high-volume, latency-sensitive workloads, now available…
- Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy
Meta came out with a banger paper https://arxiv.org/pdf/2606.00206, but it did not look at various quantizations supported in llama.cpp. So I did a run on 50…
- Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
AI Engineering from Scratch is an MIT-licensed curriculum: 523 lessons across 20 phases, from linear algebra and backprop to transformers, LLMs, agents, and…
- Are there machine learning subfields that are becoming irrelevant (or is irrelevant)? [D]
I was reading a paper that surveyed the field of neural architecture search, where it said within 5 years, around 3000+ new models were proposed. The amount…
- GPT-3 is discontinued today
It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our…
- Another "Harness matters" post (codex cli > pi and opencode)
I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my 2x3090 +…
- autotrust/JEV-27B so far best open-source Jev model
Six public decision benchmarks, one protocol: JEV-27B averages 84.07, TypeSafe Jev 1.13 83.85 Close to Jev at the level of whole probability distributions:…
- I built a LinkedIn AI slop detector using JEV
Built using Claude Code as a local Chrome extension, it scans the text and rates it based on a number of questions. The box in the corner gives you a summary…
- DeepSeek has 64% market share today 3 months ago, it had at 25% open models are winning https://t.co/ROsouLx1BZ
DeepSeek has 64% market share today 3 months ago, it had at 25% open models are winning https://t.co/ROsouLx1BZ
- Trying something new - introducing CyberPVP - in other words CyberKimi vs other AI models competing to solve complex…
Trying something new - introducing CyberPVP - in other words CyberKimi vs other AI models competing to solve complex cyber tasks. you can see it live here:…
官方发布
- Introducing Grok 4.7
- Introducing GPT-6 Sol and Luna
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
- codex 0.158.0
New Features Configure copy-on-select and right-click paste in the fullscreen TUI. Copied transcript selections now preserve Markdown formatting. (#47639,…
- Coding sessions are longer and use more context. Claude Opus 5.5 is built with that in mind.
- Better prompt caching for GPT-6
Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
- Gemini 3.8 text-to-speech says hello
Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.
- Yes, Claude can do Nine Loops
Guest writer and physicist Matt von Hippel shares what happened when he issued a challenge to AI companies to solve a problem in his former subfield of…
- Claude discovers a novel enzyme system with CRISPR-like repeats
We’re announcing a new life sciences research group and laboratory at Anthropic. This post introduces the team behind this work and shares early results in…
- Turn your REST APIs into MCP tools with Google Cloud API Gateway
Google Cloud API Gateway now acts as a native remote Model Context Protocol (MCP) server, eliminating the need to build and maintain custom middleware to…
- Introducing Support for Local AI Models in the Antigravity SDK
The Google Antigravity SDK now empowers developers to execute offline, agentic workflows locally using models like Gemma 4 26B A4B via LiteRT. This update…