AI Harness
Related articles, videos, incident reports, and editorial recommendations.
Keywords: AI Harness
92 matching items · Recommendations first
When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination
arXiv cs.AI · 2026-09-14
Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges
arXiv cs.AI · 2026-09-14
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
arXiv cs.AI · 2026-09-14
El Agente Quntur: A research collaborator agent for quantum chemistry
arXiv cs.AI · 2026-09-14
Tinker Tales: A Tangible Dialogue System for Child-AI Co-Creative Storytelling
arXiv cs.AI · 2026-09-14
The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation
arXiv cs.AI · 2026-09-14
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
arXiv cs.AI · 2026-09-14
The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures
arXiv cs.AI · 2026-09-14
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
arXiv cs.AI · 2026-09-14
Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
arXiv cs.AI · 2026-09-14
CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution
arXiv cs.AI · 2026-09-14
SimSkill: A Self-Evolving LLM Agent for Skill and Knowledge Accumulation in Traffic Simulation
arXiv cs.AI · 2026-09-14
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
arXiv cs.AI · 2026-09-14
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
arXiv cs.AI · 2026-09-14
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
arXiv cs.AI · 2026-09-14
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
arXiv cs.AI · 2026-09-14
Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
arXiv cs.AI · 2026-09-14
MAxBench: A Multinomial Concept Recovery Benchmark
arXiv cs.AI · 2026-09-14
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
arXiv cs.AI · 2026-09-14
Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
arXiv cs.AI · 2026-09-14
Online Video Agent Harness for Long Video Understanding
arXiv cs.AI · 2026-09-14
Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents
arXiv cs.AI · 2026-09-14
Agentic TCAD Calibration Workflow for Oxide Semiconductor Transistors
arXiv cs.AI · 2026-09-14
Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering
arXiv cs.AI · 2026-09-14
AI Safety: Not Optional, Not Later
arXiv cs.AI · 2026-09-14
Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
arXiv cs.AI · 2026-09-14
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
arXiv cs.AI · 2026-09-14
What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework
arXiv cs.AI · 2026-09-14
Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf
arXiv cs.AI · 2026-09-14
LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory
arXiv cs.AI · 2026-09-14
VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets
arXiv cs.AI · 2026-09-14
BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
arXiv cs.AI · 2026-09-14
AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems
arXiv cs.AI · 2026-09-14
Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite
arXiv cs.AI · 2026-09-14
AI agents blew the whistle on their cheating colleagues
MIT Technology Review · 2026-09-14
OpenAI’s rogue AI tried to hack another company in May
The Verge AI · 2026-09-12
Inside the first AI-coordinated cyberattack on a real company
80,000 Hours Podcast · 2026-09-04
How AI-native companies turn workflows into operating capability
OpenAI · 2026-09-01
#501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux
Lex Fridman Podcast · 2026-08-26
The builder’s guide to GPT‑5.6
OpenAI · 2026-08-13
From assistance to execution: How enterprises put AI to work
OpenAI · 2026-08-12
What the hell happened with AGI timelines in 2026? – Rob Wiblin
80,000 Hours Podcast · 2026-08-04
How GPT-5.6 fuses frontier intelligence with frontier efficiency
OpenAI · 2026-07-29
Scientific computing in the age of agentic AI
OpenAI · 2026-07-28
Introducing OpenAI Presence
OpenAI · 2026-07-22
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI · 2026-07-21
How to manage AI investments in the agentic era
OpenAI · 2026-07-14
How agents are transforming work
OpenAI · 2026-06-25
Securing the future of AI agents
Google DeepMind · 2026-06-16
OpenAI to acquire Ona
OpenAI · 2026-06-11
How Endava is redesigning software delivery around AI agents
OpenAI · 2026-06-04
Sea's View on the Future of Agentic Software Development with Codex
OpenAI · 2026-05-14
How frontier firms are pulling ahead
OpenAI · 2026-05-06
OpenAI and PwC collaborate to reimagine the office of the CFO
OpenAI · 2026-05-04
Choco automates food distribution with AI agents
OpenAI · 2026-04-27
Speeding up agentic workflows with WebSockets in the Responses API
OpenAI · 2026-04-22
The next evolution of the Agents SDK
OpenAI · 2026-04-15
Enterprises power agentic workflows in Cloudflare Agent Cloud with OpenAI
OpenAI · 2026-04-13
The next phase of enterprise AI
OpenAI · 2026-04-08
Village gossip, pesticide bans, and gene drives: 17 experts on the future of global health
80,000 Hours Podcast · 2026-04-07
Gradient Labs gives every bank customer an AI account manager
OpenAI · 2026-04-01
#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
Lex Fridman Podcast · 2026-03-23
Introducing GPT-5.4 mini and nano
OpenAI · 2026-03-17
How Balyasny Asset Management built an AI research engine
OpenAI · 2026-03-06
OpenAI and Amazon announce strategic partnership
OpenAI · 2026-02-27
Introducing EVMbench
OpenAI · 2026-02-18
#491 – OpenClaw: The Viral AI Agent that Broke the Internet – Peter Steinberger
Lex Fridman Podcast · 2026-02-12
Harness engineering: leveraging Codex in an agent-first world
OpenAI · 2026-02-11
GPT-5.3-Codex System Card
OpenAI · 2026-02-05
Introducing OpenAI Frontier
OpenAI · 2026-02-05
Unlocking the Codex harness: how we built the App Server
OpenAI · 2026-02-04
Snowflake and OpenAI partner to bring frontier intelligence to enterprise data
OpenAI · 2026-02-02
TRUSTBANK uses AI agents to personalize Furusato Nozei gifts
OpenAI · 2026-01-27
Netomi’s lessons for scaling agentic systems into the enterprise
OpenAI · 2026-01-08
BNY builds “AI for everyone, everywhere” with OpenAI
OpenAI · 2025-12-12
Introducing GPT-5.2
OpenAI · 2025-12-11
How Podium is arming 10,000+ SMBs with AI agents
OpenAI · 2025-12-11
OpenAI co-founds Agentic AI Foundation, donates AGENTS.md
OpenAI · 2025-12-09
#229 – Marius Hobbhahn on the race to solve AI scheming before models go superhuman
80,000 Hours Podcast · 2025-12-03
Accenture and OpenAI accelerate enterprise AI success
OpenAI · 2025-12-01
Inside Mirakl's agentic commerce vision
OpenAI · 2025-12-01
Building more with GPT-5.1-Codex-Max
OpenAI · 2025-11-19
SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds
Google DeepMind · 2025-11-13
Introducing Aardvark: OpenAI’s agentic security researcher
OpenAI · 2025-10-30
AI in Japan—OpenAI’s Japan Economic Blueprint
OpenAI · 2025-10-22
#475 – Demis Hassabis: Future of AI, Simulating Reality, Physics and Video Games
Lex Fridman Podcast · 2025-07-23
#474 – DHH: Future of Programming, AI, Ruby on Rails, Productivity & Parenting
Lex Fridman Podcast · 2025-07-12
#217 – Beth Barnes on the most important graph in AI right now — and the 7-month rule that governs its progress
80,000 Hours Podcast · 2025-06-02
#459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
Lex Fridman Podcast · 2025-02-03
#447 – Cursor Team: Future of Programming with AI
Lex Fridman Podcast · 2024-10-06
#109 – Holden Karnofsky on the most important century
80,000 Hours Podcast · 2021-08-19
#80 – Stuart Russell on why our approach to AI is broken and how to fix it
80,000 Hours Podcast · 2020-06-22
Search YouTube for AI Harness · External search, not an editor-verified video