Skip to content

Normal Science

Brain healing
GraphAuthors

A reading list for frontier science

Articles across AI, biotech, forecasting, and emerging tech.

Recommendation GraphExplore who recommends whom across the networkBrowse AuthorsProfiles, influences, and key works
Weekly Digest — Free
Join researchers, founders, and analysts · Unsubscribe anytime

Categories

AllAIForecastingBioTechMetascienceSecurity / OSINTAI SafetyFinanceManufacturingEnergyCryptoStartups

Time

Sort

Today

esengine/DeepSeek-Reasonix

·GitHub Trending·20m ago

DeepSeek-native AI coding agent for your terminal. Engineered around prefix-cache stability — leave it running. English · 简体中文 · Guide · ACP · Extensions · Spec · Website · Discord A DeepSeek-native AI coding agent for your terminal. A config- and plugin-driven harness — a single static Go binary, tuned around DeepSeek's prefix cache so token costs stay low across long sessions. Important Community · 加入社区 — bilingual Discord for setup help (#help / #求助), workflow showcases, and feature ideas. → ...

firecrawl/pdf-inspector

·GitHub Trending·20m ago

Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.pdf-inspector Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for Python, Node.js, and browser WebAssembly. Built by Firecrawl to handle text-based PDFs locally in under 200...

huangruiteng/loopx

·GitHub Trending·20m ago

Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executable todos, evidence logs, and verifiable handoffs. LoopX The local control plane for long-running AI agent work. Keep objectives, gates, todos, evidence, quota, and handoffs stable while Codex, Claude Code, Cursor, or your own runtime executes bounded turns. Public website · Docs · Try LoopX · See real...

uber/ADR

·GitHub Trending·3h ago

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.ADR: Agentic AI Detection and Response ADR (Agentic AI Detection and Response) is an enterprise security system for AI agents. It helps organizations secure employee-facing agents such as Cursor, Claude Code, and Codex, as well as customer-facing agents such as AI support agents. ADR is deployed in production at Uber, and the accompanying paper was accepted to MLSys 2026: Paper P...

Yesterday

WeatherNext: AI model achieves breakthrough in forecasting cyclones

·DeepMind·16h ago

DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness

·Apple ML Research·1d ago

Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such as “Which actor from the film Heat won at least one Academy Award?”, which requires (1) distinguishing between multiple films sharing the same title and (2) reasoning across a large set of actors to gather and integrate evidence. Existing QA benchmarks rarely evaluate both challenges jointly. To addr...

This Week

livekit/agents

·GitHub Trending·1d ago

A framework for building realtime voice AI agents 🤖🎙️📹 Looking for the JS/TS library? Check out AgentsJS What is Agents? The Agent Framework is designed for building realtime, programmable participants that run on servers. Use it to create conversational, multi-modal voice agents that can see, hear, and understand. Features Flexible integrations: A comprehensive ecosystem to mix and match the right STT, LLM, TTS, and Realtime API to suit your use case. Integrated job scheduling: Built-in task...

Assessment of open AI math results

paulpauper·5d ago4pts

It's hard for an ordinary person to understand the complexity of these tasks. I'm no mathematician, and I don't see a difference between e.g., results 3 and 10. So I had GPT-5.6 Sol Pro and Fable 5 Max classify these using @EpochAIResearch OpenMath's rubric: — "Solid Result": A

shiyu-coder/Kronos

·GitHub Trending·2d ago

Kronos: A Foundation Model for the Language of Financial Markets Kronos: A Foundation Model for the Language of Financial Markets Deutsch | Español | Français | 日本語 | 한국어 | Português | Русский | 中文 Kronos is the first open-source foundation model for financial candlesticks (K-lines), trained on data from over 45 global exchanges. 📰 News 🚩 [2025.11.10] Kronos has been accpeted by AAAI 2026. 🚩 [2025.08.17] We have released the scripts for fine-tuning! Check them out to adapt Kronos to your own ...

antirez/ds4

·GitHub Trending·2d ago

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm DwarfStar is a small native inference engine optimized first for DeepSeek V4 Flash. It also supports GLM 5.2 and, on very high-memory machines, DeepSeek V4 PRO. It is self-contained and deliberately narrow, not a general GGUF runner. Model loading, prompt rendering, tool calls, KV state, the HTTP server, and the coding agent are built and tested together. The repository also includes tools and data for GGUF, imatrix, qualit...

Taming Outlier Tokens in Diffusion Transformers

·Apple ML Research·2d ago

We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models remains underexplored. We show that this phenomenon appears in both the encoder and denoiser of modern Representation Autoencoder (RAE)-DiT pipelines: pretrained ViT encoders can produce outlier represen...

Does Forecasting Have Room At The Top?

Scott Alexander·Astral Codex Ten·3d ago

Superforecasting is the art/science/sport of predicting the future - for example, who will win elections, which countries will fight wars, when key technologies will be discovered. Over the past few years, it went from an obscure academic subfield to a multibillion dollar industry in the form of prediction markets. More recently, AI superforecasters have come close to the accuracy of top humans, and their performance is rising rapidly. In a year or two, we’ll see one of the following patterns:Ei...

An Algorithmic Failure Beneath the Secret Ballot

Lydia Owens·Freedom to Tinker·3d ago

Authored by Max Springer The secret ballot is one of the load-bearing walls of democracy. In Georgia, as in most states, this allows you to know whether your neighbor voted, but never who they voted for. This simple guarantee allows for transparent elections while deterring voter intimidation. Yet, in multiple states including Georgia, that guarantee is actively at risk. To keep the published record anonymous, ballot scanners shuffle the electronic records randomly before releasing them. But a s...

Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity

Jack Clark·Import AI·3d ago

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.Subscribe nowSelf-sustaining and self-replicating AI viruses are here:…Open weight LLMs + a well-designed harness = a persistent, self-sufficient virus…AI researchers have built a prototype computer virus which uses AI models to compromise computers, then uses their underlying GPU resources to run inference, letting it smartly figu...

bytedance/deer-flow

·GitHub Trending·4d ago

An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.🦌 DeerFlow - 2.0 English | 中文 | 日本語 | Français | Русский On February 28th, 2026, DeerFlow claimed the 🏆 #1 spot on GitHub Trending following the launch of version 2. Thanks a million to our incredible community — you made this happen! 💪🔥 DeerFlow (Deep Explor...

huggingface/speech-to-speech

·GitHub Trending·4d ago

Build local voice agents with open-source models Speech To Speech: Build voice agents with open-source models A low-latency, fully modular voice-agent pipeline: VAD -> STT -> LLM -> TTS, exposed through an OpenAI Realtime-compatible WebSocket API. Every component is swappable. The LLM slot speaks OpenAI-compatible protocols, so you can point it at a hosted provider, at HF Inference Providers, or at a vLLM or llama.cpp server on your own hardware for a fully local, fully open stack. This pipeline...

paperswithbacktest/awesome-systematic-trading

·GitHub Trending·4d ago

A curated list of awesome libraries, packages, strategies, books, blogs, tutorials for systematic trading. Awesome Systematic Trading 希望阅读中文版?点我 We are collecting a list of resources papers, softwares, books, articles for finding, developing, and running systematic trading (quantitative trading) strategies. What will you find here? 97 libraries and packages for research and live trading 40+ strategies described by institutionals and academics 55 books for beginners and professionals 23 videos an...

microsoft/TRELLIS.2

·GitHub Trending·4d ago

Native and Compact Structured Latents for 3D Generation Native and Compact Structured Latents for 3D Generation https://github.com/user-attachments/assets/63b43a7e-acc7-4c81-a900-6da450527d8f (Compressed version due to GitHub size limits. See the full-quality video on our project page!) TRELLIS.2 is a state-of-the-art large 3D generative model (4B parameters) designed for high-fidelity image-to-3D generation. It leverages a novel "field-free" sparse voxel structure termed O-Voxel to reconstruct ...

50 - Eli Lifland on AI 2027

Daniel Filan·AXRP (Daniel Filan)·4d ago

YouTube link Remember AI 2027? Not AI 2040, the newest coolest thing AI Futures Project has done, but AI 2027, their OG product? At long last, we have an AXRP episode about it. Enjoy! Topics we discuss: What is AI 2027? What happens in AI 2027? Why two endings? Who did what? Why superhuman AI in 2027? Forecasting time horizon growth When do time horizons go infinite? Forecasting effective compute growth From superhuman coders to superintelligence How many AI companies? What AGI will want What mi...

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

·Apple ML Research·4d ago

Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar to language models, MLLMs for image understanding tasks encounter challenges like hallucination. In MLLMs, hallucination can occur not only by stating incorrect facts but also by producing responses that are inconsistent with the image content. A primary objective of alignment for ...

Older

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

donsupreme·3mo ago470pts

Researchers say results mark a really ‘profound change in technology that will reshape medicine’

Accelerating Gemma 4: faster inference with multi-token prediction drafters

amrrs·3mo ago640pts

An overview of how Multi-Token Prediction (MTP) drafters are making Gemma 4 models up to 3x faster at inference.

A couple million lines of Haskell: Production engineering at Mercury

unignorant·3mo ago399pts

What it takes to run 2 million lines of Haskell in production at a fintech company serving 300,000 businesses.

Using “underdrawings” for accurate text and numbers

samcollins·3mo ago359pts

A technique for accurate text and numbers in AI-generated images: generate the layout deterministically, then ask the image model to paint on top.

ProgramBench: Can language models rebuild programs from scratch?

jonbaer·3mo ago129pts

Abstract page for arXiv paper 2605.03546: ProgramBench: Can Language Models Rebuild Programs From Scratch?

ZAYA1-8B matches DeepSeek-R1 on math with less than 1B active parameters

steveharing1·3mo ago87pts

Who should care If you work with math, science problems, or complex coding tasks and you're looking for something small enough to run locally or cheaply via API, this is worth serious evaluation. The benchmark numbers at 760M active parameters are not normal and the Markovian RSA boost means performance scales with compute budget rather than hitting a fixed ceiling. If you're building agent workflows that need reliable tool calling or multi-step instruction following, look elsewhere fo

Show HN: Apple's SHARP running in the browser via ONNX runtime web

bring-shrubbery·3mo ago170pts

Hi HN, author here. SHARP is Apple's recent single-image 3D Gaussian splatting model (https://arxiv.org/abs/2512.10685). Their reference code is PyTorch + a pretty heavy pipeline; I wanted to see if it could run in a browser with no server hop, so I exported the predictor to ONNX and ran it via onnxruntime-web with the WebGPU EP.What works: drop in an image, get a .ply you can download or preview live, all on your machine — your image never leaves the tab. The model is large (~2.4 GB sidecar) so first load is slow on a cold cache, but inference itself is a few seconds on a recent Mac.Caveats: SHARP's released weights are research-use only (Apple's model license, not the code's). I host the exported ONNX on R2 so thedemo "just works", but you can also export your own from the upstream Apple repo and upload locally.Happy to talk about it in the comments :)

Text-to-CAD

softservo·3mo ago146pts

An open source harness for generating CAD models. Contribute to earthtojake/text-to-cad development by creating an account on GitHub.

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

gmays·3mo ago149pts

Abstract page for arXiv paper 2604.26752: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

Learning the Integral of a Diffusion Model

benanne·3mo ago140pts

A deep dive on flow maps.

The Road to a Billion-Token Context

pseudolus·3mo ago38pts

Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

maxloh·14d ago19pts

Which LLM models write alike? A heat map built from their words alone.

Transformers Are Inherently Succinct (2025)

bearseascape·3mo ago45pts

Abstract page for arXiv paper 2510.19315: Transformers are Inherently Succinct

J-space comparisons across open models

babelfish·22d ago17pts

Six experiments extending the Verbalizable-Workspace paper on open models: temporal reach, training dynamics, cross-model transplant, capacity scaling, corpus dependence, MoE.

Anthropic/OpenAI may be spending more than $1000 for every $100 you pay them

gctwnl·2mo ago14pts

Coding with LLMs (Claude Code, OpenAI Codex) is often presented as the ‘killer app’ for Generative AI. But looking at data, it seems the one piece of the puzzle missing is actual cost. …

Show HN: Adam – An embeddable cross-platform AI agent library

marcobambini·3mo ago18pts

An embeddable cross-platform AI agent library written in C. Cloud and local LLMs, tool calling, long-term memory, voice, sessions, research mode, self-evolving loops. The SQLite of agent frameworks: small, portable, just works. - sqliteai/adam

Show HN: I trained a language model that thinks the capital of Japan is Paris

farisallafi·1mo ago13pts

DIMBA II: a 288M-parameter masked-diffusion language model on a bidirectional Mamba-2 backbone. What worked, what failed, and what we

Nobody checked which state IBM's flagship quantum chemistry results compute

purestatelabs·8d ago8pts

This is the data-and-code archive for the preprint "Which state are you converging? A spin audit of sample-based quantum diagonalization benchmarks on iron–sulfur clusters." Sample-based quantum diagonalization (SQD, also called QSCI) selects electronic configurations from quantum-circuit samples and diagonalizes the Hamiltonian in that subspace. Its iron–sulfur demonstrations, on the [2Fe-2S] and [4Fe-4S] clusters, are among the headline evidence for utility-scale quantum chemistry. Thi

Grok Build 0.1: Intelligence, Performance and Price Analysis

himata4113·1mo ago11pts

Analysis of xAI's Grok Build 0.1 0616 and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.

What I Learned from Reimplementing 40 Multi-Agent LLM Papers

syumei·21d ago11pts

Kimi K3 Intelligence, Performance and Price Analysis

theanonymousone·21d ago10pts

Analysis of Kimi's Kimi K3 and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.

Talk Is Cheap: The Operational Impact of LLM Use

oudlys·2mo ago9pts

What the data says about the operational impacts of LLM use in the software industry

Show HN: I put a $2.43 necklace on 3 outfits. VLMs priced it at $19 to $104

BrianneLee011·9d ago5pts

Context-driven valuation bias and halo effects across six multimodal LLMs (companion study to Lee, 2026) - BraveAnn011/ai-halo-valuation-bias

GLM-5.2: The Most Powerful Open Model yet and the Brutal Reality of Running It

ermantrout·1mo ago9pts

Z.ai’s GLM-5.2 is the new #1 open-weight model — 753B params, MIT license, a 1M-token context and a real architecture trick (IndexShare). But the weights are 1.51TB. What owners and the benchmarks actually say, and the honest hardware reality of running it at home.

The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Anon84·1mo ago4pts

Abstract page for arXiv paper 2309.12288: The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Show HN: I benchmarked LLM agents on fixing real-world security vulnerabilities

ggattip·2mo ago4pts

I built a benchmark with 20 real CVEs across 18 Python projects (Pillow, GitPython, yt-dlp, urllib3, etc). I've run it over 5 LLM agents (3 OpenAI, 2 poolside) and 3 different prompts (full advisory, locate, diagnose) with a total of 300 runs. The agents are tasked to fix security vulnerabilities in a sandboxed environment and they are scored against a hidden security tests from the maintainer's own fix.Best solve rate was 50%. On the other 50%, some fixes are sometimes coherent and pass all regression tests, but vulnerability still present.The main differentiator I found between models is cost: gpt-5.5 at 12× more expensive than gpt-5.4-mini while producing statistically similar results. Within-family performance gaps are small, which points out the difference is likely due to model training data. I also did a power analysis and the task count needed to detect a meaningful within-family edge at ~700.Full write-up: https://giovannigatti.github.io/cve-benchCode: https://github.com/GiovanniGatti/cve-bench

Following the Text Gradient at Scale

bearseascape·3mo ago5pts

RL Throws Away Almost Everything Evaluators Have to Say

What Are Tokens in LLMs?

s1monb·2mo ago7pts

How LLMs split text into tokens, the BPE algorithm, and why

Drone-Bench: Tracking simple drone surveillance capabilities of frontier models

lukaspetersson·13d ago4pts

We’re releasing Drone-Bench, a benchmark measuring how well AI models can write code to surveil real-world environments on low-cost drone hardware.

I trained a 113M-parameter earthquake LLM from absolute scratch

jzsfg·25d ago5pts

Contribute to jiazhe868/nanogpt-seis development by creating an account on GitHub.