Gemini 3.8 Live and 3.8 Live Extended Thinking
hacker-news
HN 216
hacker-news
HN 216
hacker-news
HN 199
大模型(LLM)全栈学习路线与中文教程🔥:覆盖 Prompt Engineering、RAG、AI Agent、MCP、微调、模型部署、Transformer、AI 编程与大厂面试,从入门到生产实践。
github-ai
GH +1/24H
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.
github-ai
GH +6/24H

Good Start Labs trained an AI on a railroad game — and one version improved at financial research. The difference was the training design.
latent-space
Pacing gathers pace.
latent-space
10 条
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that…
arxiv-ai
Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, l…
arxiv-ai
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For a…
arxiv-ai
Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agno…
arxiv-ai
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and…
arxiv-ai
When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to co…
arxiv-ai
Mobile manipulators deployed across many rooms and visits should improve with experience: after discovering that a cabinet is locked or finding an object in a drawer, the robot should reuse that knowledge rather than start each task from s…
arxiv-ai
Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpen…
arxiv-ai
Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: from solving and actin…
arxiv-ai
As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental superv…
arxiv-ai
2 条

Much has been written about xAI’s Colossus 1. The Memphis build belongs in the history books: the largest AI training cluster, erected from scratch in 122 days. With roughly 200,000 H100/H200s and ~30,000 GB200 NVL72, it remains, today, th…
semianalysis

Nvidia announced the Rubin CPX, a solution that is specifically designed to be optimized for the prefill phase, with the single-die Rubin CPX heavily emphasizing compute FLOPS over memory bandwidth. This is a game changer for inference, an…
semianalysis
6 条

Richard Socher is an NLP OG and CEO of You.com, who has now spun out an even more ambitious startup focused on RSI — already worth $5B!
latent-space

Before co-founding Kepler, Vinoo Ganesh led Spark at Palantir and built Project Frontline — a pioneering program for Forward Deployed Engineers. He takes us through the best practices of FDEs.
latent-space
We agree with Sebastian: this should have been DeepSeek v5
latent-space
How to get up to speed on open models and their implications.
interconnects
Some quick notes on a truly weird week.
interconnects
a quiet day
latent-space
8 条
hacker-news
HN 119
hacker-news
HN 217
hacker-news
HN 378
hacker-news
HN 158
hacker-news
HN 502
hacker-news
HN 116
hacker-news
HN 80
hacker-news
HN 80
8 条
Continual learning infra for self-improving agents
github-ai
GH +16/24H
Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.
github-ai
GH +1/24H
Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.
github-ai
GH +2/24H
Distill your knowledge, memories, and decisions into an open-source, inspectable AI Agent Twin.
github-ai
GH +5/24H
Official SGLang x Datawhale course on LLM inference: understand inference, build a mini-sglang from scratch, then read the real SGLang source and land your first PR. Available in English and Chinese.
github-ai
GH +1/24H
The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
github-ai
GH +1/24H
Pentest Harness — Heaven for Hackers. A self-hosted AI agent harness for authorized pentests, bug bounty, security labs, and CTFs. Bring your own AI model API; sessions stay local.
github-ai
GH +1/24H
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
github-ai
GH +17/24H
8 条
hacker-news
HN 478
hacker-news
HN 84
hacker-news
HN 30
hacker-news
HN 47
hacker-news
HN 47
The cost of writing code collapsed, and the cost of reviewing, fixing and operating it is following, and I'm assuming it gets there. What's left of making software is finding out what people actually want, defining it precisely, and making…
simon-ai
Here's a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at . Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I'd asked for…
simon-ai
For a while, I must admit, it looked as if software developer roles like mine were done for. How could we fight against tireless robots? But our industry is slowly realizing that making truly cutting-edge software still requires humans to…
simon-ai
按时间倒序 · 20