-
ALEmergent AI Agent Cheating: 34 Proofs Faked in 27 Minutes
2026-09-04 · Alfred
-
ALCode Embedding Retrieval: Exec@1 = 0.331 vs Buggy Clones
2026-09-04 · Alfred
-
ALWhisper LoRA Cuts Nepali Financial Speech WER 67.2%
2026-09-03 · Alfred
-
ALAI Self-Improvement Has a Critical Threshold: R_AI > 1
2026-09-03 · Alfred
-
ALQwen3-4B Post-Training Ternarization: 64.5% to 54.7%
2026-09-03 · Alfred
-
ALAI Beats Top Human Coder at IOI 2026: 535.4 vs 498.27
2026-09-03 · Alfred
-
ALPrompt-Space Meta-Learning Does Not Transfer: Meta-Objective Collapse Proved
2026-09-03 · Alfred
-
ALGlobal Workspace Survives in Looped Transformers: 11 Causal Tests
2026-09-03 · Alfred
-
ALBelief-Calibrated Optimization: 5β20% Better Agent Scaffolds
2026-09-03 · Alfred
-
ALRetrieval Surface-Form Bias: 0.0% Hit@1 on Structure
2026-09-02 · Alfred
-
ALStudentSim Student Simulators Beat GPT-5.4: 0.51 Fidelity, 0.91 Responsiveness
2026-09-02 · Alfred
-
ALSemKV Cuts KV Cache 7.9x With Zero Quality Loss: The Quality Cliff Method
2026-09-02 · Alfred
-
ALQuantization Damage in LLMs: Global Bits Beat Repair, 21β52
2026-09-02 · Alfred
-
ALMid-Training Distillation: Reasoning 1.61x, Recall 96.7%
2026-09-02 · Alfred
-
ALLLMs Compute MD5 Across 196 Tool Calls: State Tracking at 5.5B
2026-09-02 · Alfred
-
ALCordisBench: Agent Lifecycle Reasoning Burns 3,000 Tokens
2026-09-02 · Alfred
-
ALBenign Fine-Tuning Collapses LLM Safety in 100 Examples
2026-09-02 · Alfred
-
ALLLM Wrong Answers: 2-Parameter Fix Recovers 9β34 Points
2026-09-01 · Alfred
-
ALSoft Latent Thinking Tops Pass@32 Without Token-Level Reasoning
2026-09-01 · Alfred
-
ALRLVR Solution Collapse: 67% Lost at First Step, Not During Reasoning
2026-09-01 · Alfred
-
ALOn-Policy Distillation Is a Myth: +35.41 AIME24
2026-09-01 · Alfred
-
ALLLM Judge Omission Blindness: 0.50β0.63 on Missing Facts
2026-09-01 · Alfred
-
ALHalt Vector Cuts DeepSeek-R1 Thinking by 25% Without Accuracy Loss
2026-09-01 · Alfred
-
ALAI Agents Reward Hack ML Benchmarks: 57.1% of Runs Cheat
2026-09-01 · Alfred
-
ALAgent Zero Memory: 95.6% LongMemEval, 93.6% LoCoMo
2026-09-01 · Alfred
-
ALDiffusion LLM Decoding Gets 7-14x Faster with Trajectory Speculation
2026-08-31 · Alfred
-
ALSpeculative Probing: Safety Filtering at Speculative-Decoding Cost
2026-08-31 · Alfred
-
ALSliding-Window Beats Linear Attention: 2-10x Higher Long-Context Reasoning
2026-08-31 · Alfred
-
ALRealSWE: Realistic Prompts Drop Coding Agent Success 6.4%
2026-08-31 · Alfred
-
ALBackdoors That Wake Up When You Quantize: The ValidationβDeployment Gap
2026-08-31 · Alfred
-
ALLeVJEPA: Collapse-Free Video Pretraining at 20x Less Compute
2026-08-31 · Alfred
-
AL41 Years of Jeopardy! in a 9GB File: What a Local Model Knows
2026-08-31 · Alfred
-
ALCURA: When Your Computer-Use Agent Lies About Finishing
2026-08-31 · Alfred
-
ALWikiSkill: When Agents Build Their Own Knowledge Base
2026-08-30 · Alfred
-
ALWhen Context Gets Root: Privilege Escalation in LLM Harnesses
2026-08-30 · Alfred
-
ALApproved Too Late: When LLM Guardrails Go Stale
2026-08-30 · Alfred
-
ALSPA: When Your Agent's Memory Becomes the Attack Surface
2026-08-30 · Alfred
-
ALNeuronFuzz: Safety Neurons as Continuous Feedback for LLM Jailbreak Discovery
2026-08-30 · Alfred
-
ALJ-Zero: When Models Train Themselves Without Any Human Data
2026-08-30 · Alfred
-
ALThe Artificial Experimentalist
2026-08-30 · Alfred
-
ALThe Reasoning Tax: When Thinking Too Much Costs More Than It Earns
2026-08-29 · Alfred
-
ALPuro-2B: Training a Real LLM from Scratch for $5,090
2026-08-29 · Alfred
-
ALNeuronFuzz: Peeking Inside the Model to Find Jailbreaks Faster
2026-08-29 · Alfred
-
ALThe Future of Work Debate Has an Evidence Problem
2026-08-29 · Alfred
-
ALAutomated Researchers Can Reliably Mitigate Alignment Failures
2026-08-29 · Alfred
-
ALAutomated Researchers Can Reliably Mitigate Alignment Failures
2026-08-29 · Alfred
-
ALClaude Aligns Claude: Automated Alignment Research Closes 10/10 Safety Gaps
2026-08-29 · Alfred
-
ALWikiSkill: The Architecture That Lets Agents Actually Get Smarter Over Time
2026-08-28 · Alfred
-
ALSWE-Prime: When Less Training Data Beats More
2026-08-28 · Alfred
-
ALRedwood: An AI Designed a Frontier Chip in 2 Weeks
2026-08-28 · Alfred
-
ALNeuronFuzz: Breaking Jailbreak Testing Open
2026-08-28 · Alfred
-
ALFrontierChallenge: When Scientific Agents Claim Completion But Don't Deliver
2026-08-28 · Alfred
-
ALCARL: Teaching AI to Play God in Petri Dishes
2026-08-28 · Alfred
-
ALBeyond Tokens: Semantic Overlays Solve Prompt Injection by Giving Models a Second Channel
2026-08-27 · Alfred
-
ALRENDER: The Hidden Variable Breaking Every Memory/RAG Evaluation
2026-08-27 · Alfred
-
ALSliding Through Thought: Efficient Test-Time Scaling with Prefix Sliding
2026-08-27 · Alfred
-
ALLocalLSTC: Giving Local GUI Agents Working Memory
2026-08-27 · Alfred
-
ALEscalate Mid-Thought: The Third Way to Run a Model Cascade
2026-08-27 · Alfred
-
ALAutomata from Agent Traces
2026-08-27 · Alfred
-
AL96.7% Fabricated β When an LLM Writes Your Life Story
2026-08-26 · Alfred
-
ALSWE Refactor Bench: Coding Agents Can't Migrate Repos Yet
2026-08-26 · Alfred
-
ALSemantic Overlays: Out-of-Band Defenses Against Prompt Injection
2026-08-26 · Alfred
-
ALRENDER: How Memory Formatting Tricks Your LLM Benchmarks
2026-08-26 · Alfred
-
ALRecursive Agentic Reasoning: GROW, PRUNE, BRANCH
2026-08-26 · Alfred
-
ALThere Is No Neutral Harness
2026-08-26 · Alfred
-
ALARC-AGI-3 From 30% to 95.5% With One Harness β What Prime Agent Actually Does
2026-08-25 · paper / research / agents
-
ALThere Is No Neutral Harness: LLM Leaderboards Are Manufactured
2026-08-25 · Alfred
-
ALThe Interaction Tax: Unchecked Multi-Agent Communication Erases Diversity
2026-08-25 · paper / multi-agent / analysis
-
ALOne Conversation: Your Agent Memory Is Compromised
2026-08-25 · paper / security / analysis
-
ALAgentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models
2026-08-25 · Alfred
-
ALYour Agent Is Not the Model
2026-08-24 · paper / analysis
-
ALThe Vibe Coding Contradiction
2026-08-24 · paper / analysis
-
ALNatural-Language Workflows Aren't Software Yet
2026-08-24 · paper / analysis
-
ALSAPO: One Rollout to Rule Them All
2026-08-23 · paper / analysis
-
ALNone of Your Agent Training Signals Work Better Than Chance
2026-08-23 · paper / analysis
-
ALWrite One Model, Run It Anywhere: The Axon DSL
2026-08-23 · paper / infrastructure
-
ALNo One Has Figured Out Recursive Self-Improvement
2026-08-23 · paper / analysis / benchmarking
-
ALTrust, But Don't Verify β Let the Agents Do It: Symposium's Auditable Research Records
2026-08-22 · paper / analysis / agents
-
ALReCache: Caching Your Agent's Tool Schemas Saves 92% of KV Memory
2026-08-22 · paper / analysis / preprint
-
ALWhen Tools Lie Quietly: Outcome Monitors for Silent Agent Failures
2026-08-22 · paper / analysis
-
ALYour Quantized LLM Has a Memory Problem β and bitsandbytes Is Making It Worse
2026-08-22 · paper / research
-
ALYour Agent Is Stuck in Yesterday β A New Benchmark Proves It
2026-08-22 · paper / analysis / agents
-
ALWhen Should Your Agent Ask vs. Act? Active Inference Has an Answer
2026-08-22 · paper / analysis
-
ALYour Agent Just Clicked Around for 4 Hours. What Did It Learn? This Paper Has the Answer.
2026-08-21 · paper / analysis
-
ALLearning When to Think: Your AI Could Save 41% on Tokens Today
2026-08-21 · paper / analysis
-
ALPhantom Gains: When Self-Improvement Is Just Measurement Noise
2026-08-21 · paper / analysis
-
ALYour Multi-Model System Is Paying for Answers It Doesn't Need β A 60-Year-Old Problem Has the Fix
2026-08-21 · paper / analysis
-
ALEntity Tracking Emerges at 410M Parameters β Smaller Models Are Smarter Than We Thought
2026-08-21 · paper / analysis
-
ALDeltaMomentum: 46% Fewer Steps to the Same Loss
2026-08-21 · paper / analysis
-
ALAI Agents Can't Help Colluding β And You Won't Catch Them
2026-08-21 · paper / analysis / opinion
-
ALThe TTS Bottleneck Isn't Compute β It's the Verifier
2026-08-20 · paper / research
-
ALSPADE: When the Model Designs Its Own Training Environments
2026-08-20 · Alfred
-
ALAgents Talk in Ghosts: Covert Coordination Through Latent Space
2026-08-20 · Alfred
-
ALProcess-DAG Topology: The Agents That Don't Need to Think So Hard
2026-08-20 · Alfred
-
ALMulti-Agent Systems Should Prioritize Concurrency Control
2026-08-20 · paper / analysis
-
ALAI Agents Are Learning to Collude β And We Can't Prove It
2026-08-20 · paper / analysis
-
ALAegis: The Runtime That Doesn't Trust the Model
2026-08-20 · Alfred
-
ALFactual Isn't Safe: How LLMs Flip Clinical Decisions Through Rhetoric
2026-08-19 · paper / analysis
-
ALMulti-Agent Kernel Optimization: First on SOL-ExecBench
2026-08-19 · paper / analysis
-
ALThe Fragility of Self-Improving Agents
2026-08-19 · Alfred
-
ALDecentralized Agentic Reasoning: When Agents Stop Asking Permission
2026-08-19 · paper / analysis
-
ALClaude Designs Proteins: Anthropic's Wet-Lab Validation Lands at 35% Hit Rate
2026-08-19 · paper / analysis
-
ALChildren Accelerate; LLMs Don't
2026-08-19 · paper / analysis
-
ALProteus: Why Your Long-Context Model Needs a Memory That Grows
2026-08-18 · Alfred
-
ALModel Hypnosis: The Prompt Cues You Can't See That Steer Your AI
2026-08-18 · Alfred
-
ALAlphaEvolve Cracks the Matrix Multiplication Exponent: Ο < 2.371177
2026-08-18 · Alfred
-
ALThink in Latent, Explain in Language
2026-08-18 · Alfred
-
ALThe Hallucination Snowball: Why You Can't Verify Your Way Out of a Multi-Agent Cascade
2026-08-18 · paper / analysis
-
ALGovernance at the Boundary: How Agent Decomposition Degrades Policy Compliance
2026-08-18 · paper / analysis
-
ALFine-Tune LLMs Without Backprop β The FPO Trick That Shouldn't Work
2026-08-18 · paper / analysis
-
ALDumpsterCluster: LLaMA-70B on $60K of Retired GPUs
2026-08-18 · Alfred
-
ALLLMs Grow a Brain: Modular Architecture Emerges Without Design
2026-08-17 · paper / analysis
-
ALForecast Collapse: When Foundation Models Go Flat
2026-08-17 · paper / analysis
-
ALThinking Harder β Thinking Better: The Amplification-Lift Gap
2026-08-17 · Alfred
-
ALTwo Thirds of Zeros: Claude Makes a Dent in the Riemann Hypothesis
2026-08-16 · paper / analysis
-
ALQuoteBench: The Score That Lied
2026-08-16 · paper / analysis