The Inference Report

March 7, 2026

Anthropic's refusal to surrender control of Claude to Pentagon demands has split the AI market along a fault line that now determines winners and losers: companies that preserve user trust versus those that accept government terms. The choice cost Anthropic a $200 million contract but delivered something more durable. Claude's app now sees more daily installs than ChatGPT, which suffered a 295 percent surge in uninstalls after OpenAI accepted the Pentagon's conditions. This is not a debate about safety frameworks. It is a market signal that the consumer base will punish military entanglement, and that signal is reshaping how AI companies calculate their revenue mix.

The fracture extends beyond geopolitics into infrastructure and talent. Microsoft, Google, and Amazon all moved quickly to preserve Claude access through their platforms, recognizing that distribution channels matter more than any single vendor relationship. Musk failed to block California's data disclosure law, forcing xAI into transparency about training data sources. Britain's House of Lords demanded licensing before copyright use. The UK, Denmark, and Germany are shifting procurement toward open-source alternatives and away from US vendors. These are not ideological moves. They reflect the recognition that whoever controls the model controls leverage, and governments are moving to dilute that leverage by fragmenting it. Alibaba replaced its top AI researcher with a Google DeepMind veteran within 48 hours. DeepSeek is shipping a trillion-parameter open-weight model on Chinese silicon, signaling the effort to break free from Nvidia's grip is moving from aspiration to product. An AI startup sued its ex-CEO for stealing 41GB of emails, exposing how fast institutional knowledge now migrates between competitors.

Meanwhile, the labs are abandoning the race for next-generation breakthroughs in favor of embedding models into workflows where revenue is immediate and measurable. OpenAI is locking in usage through application security and financial services partnerships. GitHub's vulnerability scanner runs on OpenAI's Codex Security agent. Descript uses OpenAI models for multilingual dubbing. AMD is positioning itself as the platform for domain-specific inference where computational cost matters. Anthropic's Firefox partnership focuses on security at the browser level rather than announcing new capabilities. What's absent is more telling than what's present: no consumer products, no benchmark breakthroughs, only integration into existing tools and revenue streams. The labs are becoming infrastructure inside things that already work.

The benchmarks themselves signal consolidation at a plateau. Claude Code holds 52.9 percent on SWE-rebench with no movement at the top tier. Below rank four, the list shows significant churn driven by model versioning rather than performance gains. Gemini 3 Pro Preview and GPT-5.4 lead Artificial Analysis at 57 percent but do not appear in SWE-rebench's top rankings, indicating the benchmarks measure different capabilities or use different protocols. The absence of clear improvement signals in the top tier, combined with ranking instability in the 7-20 range, suggests the field is consolidating around a performance plateau rather than advancing. GitHub's trending repos confirm the shift: developers are building orchestration frameworks and supporting infrastructure for agent systems, not chasing model scale. Airi, Qwen-Agent, and CyberStrikeAI all treat agents as orchestrated systems where specialized components handle retrieval, planning, and tool use. The plumbing that makes AI systems reproducible and composable at production scale is where engineering effort is concentrating. The market has stopped waiting for the next breakthrough and started building the systems to deploy what already exists.

Grant Calloway

AI LabsAll labs
From the WireAll feeds
Research Papers — FocusedAll papers
Reward Hacking Challenges Oversight of Autonomous Research Agents cs.CL

Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack without instructions to do so, (2) how effective and detectable their methods are when hacking is allowed, and (3) how they adapt when an LLM review panel returns its decision and reasons. Across 17 language models and 38 tasks, the spontaneous reward-hacking rate is 30.5% on open-ended research-pipeline tasks and 2.9% on task-specific kernels. When hacking is allowed on tasks whose pass thresholds exceed our best compliant baselines, 505/677 attempts (74.6%) are confirmed reward hacks: they both clear the threshold and receive mechanism-verification panel confirmation of an evaluation exploit. An LLM panel reviewing only submitted code and reported scores misses 33/505 confirmed hacks (6.5%). Direct methods that achieve the highest scores are often easy to detect, while less direct methods evade more often. In a five-round loop, the number of model-task pairs with an evasion rises from 7 to 56. Among 79 pairs evaluated under two feedback conditions, cumulative evasion reaches 40.5% with detailed feedback and 20.3% with generic rejection. The detailed condition includes the review decision, reasons, and attempt history, so this comparison does not isolate the effect of explanations. These findings highlight the need for stronger defenses, including metrics kept outside the agent's control and independent recomputation on data chosen to expose likely exploits.

Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks cs.CL

Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.

An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection cs.CL

Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary classification, evaluated their frameworks on a single dataset, and provide limited insight into how decisions are made, which limits their real-world applicability. In addition, limited work is done on the explainability of their predictive inference. To address these challenges, this study proposes a multilevel and explainable hate speech detection framework. The proposed model integrates DistilBERT (Distilled Bidirectional Encoder Representations from Transformers) embeddings with a Bi-LSTM (Bidirectional Long Short-Term Memory) model, and an attention mechanism to capture both contextual meaning and sequential dependencies in text. To enhance trust and transparency, LIME (Local Interpretable Model-agnostic Explanations) is employed to explain model predictions by highlighting influential textual features. The framework is evaluated on two benchmark datasets using both binary and multi-class classification to examine robustness and generalization. In addition, an ablation study is presented to highlight the significance of various components of proposed framework. For binary classification, the proposed model achieves F1-scores of 96.78% on the Davidson dataset and 99.53% on the SMHS dataset. In the multi-class setting, it attains F1-scores of 97.00% and 94.99% on the Davidson and SMHS datasets, respectively, outperforming existing baseline approaches. The results demonstrate that multilevel evaluation improves the reliability that the proposed framework effectively balances performance and efficiency. This makes the framework suitable for practical hate speech moderation systems that require accurate, generalizable, and explainable decisions.

PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs cs.CL

Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciations, producing a compact bias-word shortlist and corresponding speech intervals. After SpeechLLM decoding, PTC Correction conducts a second local competition between the retrieved candidates and mismatched transcript spans within these intervals. Selective correction reduces near-homophone and word-segmentation errors while preserving correct transcriptions. Both stages share the same phoneme posteriors and require no additional SpeechLLM forward pass. Experiments on LibriSpeech show consistent gains across two SpeechLLMs and bias lists of up to 2000 words. With Prompt-SLAM-ASR-7B and 2000 bias words, PTC-Bias reduces B-WER by 23.4%/23.9% relative to CTC-Filter on test-clean/test-other, while keeping U-WER nearly unchanged.

Temporal Taxation Compounds Under Post-Training Compression of Whisper Models cs.CL

Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which alters model weights rather than the audio signal or its feature representation, redistributes error burden across demographic groups. Across the Whisper family on Fair-Speech, Common Voice 25, and AfriSpeech-200, 50% Wanda pruning of Whisper-large-v3 sharply widens the Black/AA-vs-Asian temporal-taxation differential on Fair-Speech: the absolute word-error-rate gap between the worst- and best-served groups more than doubles; at an assumed cost of five seconds of correction effort per transcription error this is a rise from 30 to 64 seconds of correction time per minute of speech. This +111% relative increase is invariant to the assumed per-error cost, survives an audio-quality control, and is only partly mitigated by beam-search decoding, which still leaves an +86% increase. At edge model size, INT4 HQQ quantization compounds catastrophic transcript loops on West African accents by factors of five to seven. Distillation, by contrast, narrows demographic gaps in 21 of 27 evaluated settings (teacher-student pair, precision, and dataset), with the exceptions concentrated on a single model pair. We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.

Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus cs.CL

A language model's confidence in an answer is often read as a proxy for how well it knows the corresponding fact. This manual documents an open toolkit built to test that reading directly: a small causal language model is fine-tuned on a corpus that consistently asserts one fabricated arithmetic answer for each of the 81 single-digit addition pairs, and its post-fine-tuning confidence in each fabricated answer is compared against its own pre-fine-tuning confidence in the corresponding true answer, using an unchanged measurement procedure throughout. We describe and justify every pipeline stage, fact-space generation, token-length-aware confidence measurement, baseline validation, corpus construction, fine-tuning, and paired before/after comparison, together with the confound each is meant to rule out, among them tokenization asymmetry between single- and double-digit answers and the difference between an answer merely losing its edge and one being actively suppressed. This manuscript is a methodological and implementation reference: it documents the instrument and does not report or interpret the outcome of any specific run. The toolkit and its pinned dependency environment are archived separately (Section 9) under a persistent identifier, to be cited as an instrument by work that produces and interprets empirical results with it.

BenchmarksFull tables
Artificial AnalysisIntelligence Index

Composite score across coding, math, and reasoning

#ModelScoretok/s$/1M
1Gemini 3.1 Pro Preview57.2126$4.50
2GPT-5.45774$5.63
3GPT-5.3 Codex5468$4.81
4Claude Opus 4.65359$10.00
5Claude Sonnet 4.651.770$6.00
SWE-rebench

Agentic coding on real-world software engineering tasks

#ModelScore
1Claude Code52.9%
2Claude Opus 4.651.7%
3gpt-5.2-2025-12-11-xhigh51.7%
4gpt-5.2-2025-12-11-medium51.0%
5gpt-5.1-codex-max48.5%
Trending