AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-20 AI 日报

统计日时区:America/Los_Angeles · 已结算

BuilderIO/agent-native · FareedKhan-dev/train-llm-from-scratch · earendil-works/pi

在应用中阅读本期

开源项目

  1. BuilderIO/agent-native

    A framework for building agentic apps

  2. FareedKhan-dev/train-llm-from-scratch

    A straightforward method for training your LLM, from downloading data to generating text.

  3. earendil-works/pi

    AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI

  4. affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode,…

  5. higgsfield-ai/higgsfield

    Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters

  6. anthropics/claude-code

    Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining…

  7. virgiliojr94/book-to-skill

    Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.

  8. anthropics/financial-services

  9. coder/coder

    Secure environments for developers and their agents

  10. Tencent/BrowserSkill

    Let AI agents use your real, logged-in browser without interrupting your work. CLI + extension for browser automation across any shell-capable AI agent.

论文

  1. OpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection

    Omnidirectional video saliency prediction plays an important role in many immersive multimedia applications, including viewport-adaptive streaming and…

  2. Catena: A Comprehensive Software Suite for Large-Scale Connectomics

    The gold standard datasets for mapping connectomes are electron microscopy volumes of densely labeled neural tissue at nanometer resolution. Yet…

  3. IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

    Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many…

  4. AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

    Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances…

  5. Multi-viewpoint Geo-localization with Event Cameras

    Robot localization is an ongoing challenge that demands mapping and positioning systems that are tolerant to viewpoint change. Event cameras are attracting…

  6. Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency

    Factual hallucination is commonly defined by incorrect factual outputs. We study a paraphrase-induced hallucination setting, where a model answers a factual…

  7. Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

    Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches…

  8. How Many Humans Is a Judge Panel Worth?

    How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against empirical…

  9. GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development

    Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily…

  10. FairLMs: A Turnkey Library for Fairness in Language Models

    Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing…

行业新闻

  1. Qwen Image 2.1

  2. ChatGPT now knows what you do on other websites via ad collector

  3. AX – Google’s Open Agentic Orchestrator

  4. Pirate Face Rescues LLM Models from Deletion

  5. Chat-based Large Language Models replicate the mechanisms of a psychic's con

  6. MCP was always a bad idea?

  7. Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

    Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation.…

  8. Heretic removes restrictions from language models

  9. AI chatbots give wrong answers to financial queries 'most of the time'

  10. Show HN: A competition for small neural networks that play strategy games

社区动态

  1. Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series!

    Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both…

  2. Qwen-Image-2.1 is now supported in ComfyUI!

    Qwen-Image-2.1 is now supported in ComfyUI! Try it now and share your creations! 🖼️

  3. We’ve received so much love for Qwen-Image-2.1 over the past 24 hours, thank you!!

    We’ve received so much love for Qwen-Image-2.1 over the past 24 hours, thank you!! Also gotten a lot of questions about the license, especially around model…

  4. ZCode is now open source

    ZCode is now open source, and the reported security issues have been addressed. Source code: https://github.com/zai-org/ZCode The repo includes its desktop…

  5. Qwen-Image-2.1 released!

    Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both…

  6. Qwen Image 2.1 - official blog post

  7. Qwen-Image-2.1 weights are available on Hugging Face

  8. Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made illegal agreement on AI slowdown

  9. The famous "Car Wash" question on Jev

    But then how should we evaluate its “intelligence”? Does Jev have benchmarks comparable to the usual LLM benchmarks, or is the comparison fundamentally…

  10. I built mini Jev, a tiny open decision model for AI agents.

    Instead of generating text, it takes agent state plus available actions and directly returns action probabilities. Architecture: state: frozen Qwen3 0.6B…

官方发布

  1. Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation

    We are excited to open-source Qwen-Image-2.1, an image model in the Qwen family that balances generation quality, inference efficiency, and cost.…

  2. Qwen/Qwen-Image-2.1-PE-T2I

    text-to-image

  3. Introducing Grok Voice Transcribe 2.0

  4. Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

    Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet, built for natural conversation.

  5. Qwen3.8-LiveTranslate: Names the speaker. Carries the meaning.

    Simultaneous interpretation is not only about translating fast — it must also hear clearly and translate accurately. Qwen3.8-LiveTranslate rebuilds real-time…

  6. Partnering with Accenture on embedded evaluation

  7. zai-org/ZCode

    Z.ai's coding agent harness. Powerful, intelligent, extensible.

  8. The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface…

    The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface value claude_in_chrome), in beta for…

  9. For cache diagnostics, a response to a request that sends the cache-diagnosis-2026-04-07 beta header now always…

    For cache diagnostics, a response to a request that sends the cache-diagnosis-2026-04-07 beta header now always includes the diagnostics field. The field is…

  10. Why client SDK generation belongs in the open

    Google has partnered with Speakeasy to open-source their OpenAPI code generation suite under the AGPLv3 license, a strategic move prompted by the sudden…