2026-09-20 AI 日报
统计日时区:America/Los_Angeles · 已结算
BuilderIO/agent-native · FareedKhan-dev/train-llm-from-scratch · earendil-works/pi
开源项目
- BuilderIO/agent-native
A framework for building agentic apps
- FareedKhan-dev/train-llm-from-scratch
A straightforward method for training your LLM, from downloading data to generating text.
- earendil-works/pi
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
- affaan-m/ECC
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode,…
- higgsfield-ai/higgsfield
Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
- anthropics/claude-code
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining…
- virgiliojr94/book-to-skill
Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.
- anthropics/financial-services
- coder/coder
Secure environments for developers and their agents
- Tencent/BrowserSkill
Let AI agents use your real, logged-in browser without interrupting your work. CLI + extension for browser automation across any shell-capable AI agent.
论文
- OpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection
Omnidirectional video saliency prediction plays an important role in many immersive multimedia applications, including viewport-adaptive streaming and…
- Catena: A Comprehensive Software Suite for Large-Scale Connectomics
The gold standard datasets for mapping connectomes are electron microscopy volumes of densely labeled neural tissue at nanometer resolution. Yet…
- IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many…
- AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances…
- Multi-viewpoint Geo-localization with Event Cameras
Robot localization is an ongoing challenge that demands mapping and positioning systems that are tolerant to viewpoint change. Event cameras are attracting…
- Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency
Factual hallucination is commonly defined by incorrect factual outputs. We study a paraphrase-induced hallucination setting, where a model answers a factual…
- Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing
Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches…
- How Many Humans Is a Judge Panel Worth?
How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against empirical…
- GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development
Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily…
- FairLMs: A Turnkey Library for Fairness in Language Models
Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing…
行业新闻
- Qwen Image 2.1
- ChatGPT now knows what you do on other websites via ad collector
- AX – Google’s Open Agentic Orchestrator
- Pirate Face Rescues LLM Models from Deletion
- Chat-based Large Language Models replicate the mechanisms of a psychic's con
- MCP was always a bad idea?
- Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation.…
- Heretic removes restrictions from language models
- AI chatbots give wrong answers to financial queries 'most of the time'
- Show HN: A competition for small neural networks that play strategy games
社区动态
- Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series!
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both…
- Qwen-Image-2.1 is now supported in ComfyUI!
Qwen-Image-2.1 is now supported in ComfyUI! Try it now and share your creations! 🖼️
- We’ve received so much love for Qwen-Image-2.1 over the past 24 hours, thank you!!
We’ve received so much love for Qwen-Image-2.1 over the past 24 hours, thank you!! Also gotten a lot of questions about the license, especially around model…
- ZCode is now open source
ZCode is now open source, and the reported security issues have been addressed. Source code: https://github.com/zai-org/ZCode The repo includes its desktop…
- Qwen-Image-2.1 released!
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both…
- Qwen Image 2.1 - official blog post
- Qwen-Image-2.1 weights are available on Hugging Face
- Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made illegal agreement on AI slowdown
- The famous "Car Wash" question on Jev
But then how should we evaluate its “intelligence”? Does Jev have benchmarks comparable to the usual LLM benchmarks, or is the comparison fundamentally…
- I built mini Jev, a tiny open decision model for AI agents.
Instead of generating text, it takes agent state plus available actions and directly returns action probabilities. Architecture: state: frozen Qwen3 0.6B…
官方发布
- Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation
We are excited to open-source Qwen-Image-2.1, an image model in the Qwen family that balances generation quality, inference efficiency, and cost.…
- Qwen/Qwen-Image-2.1-PE-T2I
text-to-image
- Introducing Grok Voice Transcribe 2.0
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet, built for natural conversation.
- Qwen3.8-LiveTranslate: Names the speaker. Carries the meaning.
Simultaneous interpretation is not only about translating fast — it must also hear clearly and translate accurately. Qwen3.8-LiveTranslate rebuilds real-time…
- Partnering with Accenture on embedded evaluation
- zai-org/ZCode
Z.ai's coding agent harness. Powerful, intelligent, extensible.
- The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface…
The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface value claude_in_chrome), in beta for…
- For cache diagnostics, a response to a request that sends the cache-diagnosis-2026-04-07 beta header now always…
For cache diagnostics, a response to a request that sends the cache-diagnosis-2026-04-07 beta header now always includes the diagnostics field. The field is…
- Why client SDK generation belongs in the open
Google has partnered with Speakeasy to open-source their OpenAPI code generation suite under the AGPLv3 license, a strategic move prompted by the sudden…