arXiv · Computation and Language · 13 Aug 2026 · paper
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus on how t…
arXiv · Computation and Language · 13 Aug 2026 · paper
A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In th…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model (LLM)-based educational assistants often provide direct answers offering little incentive for students to explore or engage with course materials. We present BLADE (Bet…
arXiv · Computation and Language · 13 Aug 2026 · paper
Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuanc…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generat…
arXiv · Computation and Language · 13 Aug 2026 · paper
Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have separate embeddings in models, so occur…
arXiv · Computation and Language · 13 Aug 2026 · paper
The performance of LLM-based agents is jointly shaped by their base models and the harnesses that mediate their interaction with the environment. Because different models exhibit distinct b…
arXiv · Computation and Language · 13 Aug 2026 · paper
Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and information coverage) to support a…
arXiv · Computation and Language · 13 Aug 2026 · paper
Recent work shows superior performance when using large language models (LLMs) as formalizers instead of as end-to-end solvers for symbolic reasoning problems. Given the problem description…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction. We study a…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on man…
arXiv · Computation and Language · 13 Aug 2026 · paper
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine th…
arXiv · Computation and Language · 13 Aug 2026 · paper
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt tra…
arXiv · Computation and Language · 13 Aug 2026 · paper
Public procurement involves the allocation of substantial financial resources; therefore, continuous oversight through audits, controls, and monitoring mechanisms is essential. However, sta…
arXiv · Computation and Language · 13 Aug 2026 · paper
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before…
arXiv · Computation and Language · 13 Aug 2026 · paper
While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling…
arXiv · Computation and Language · 13 Aug 2026 · paper
Regional dialectal variation poses a fundamental challenge to natural language processing (NLP) in Bangla, where over 240 million speakers communicate across diverse regional variants that…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic uncertainty scores ca…
arXiv · Computation and Language · 13 Aug 2026 · paper
Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existi…
arXiv · Computation and Language · 13 Aug 2026 · paper
Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems re…
arXiv · Computation and Language · 13 Aug 2026 · paper
Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on Engl…
arXiv · Computation and Language · 13 Aug 2026 · paper
When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer org…
arXiv · Computation and Language · 13 Aug 2026 · paper
Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream ta…
arXiv · Computation and Language · 13 Aug 2026 · paper
Financial text is produced and interpreted within a market environment, yet financial text classifiers almost always receive text alone. We study whether financial time series are useful as…
arXiv · Computation and Language · 13 Aug 2026 · paper
Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parall…
arXiv · Computation and Language · 13 Aug 2026 · paper
As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by thes…
arXiv · Computation and Language · 13 Aug 2026 · paper
In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to de…
arXiv · Computation and Language · 13 Aug 2026 · paper
Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of a…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration. In this new era of work, it is important to understand the k…
arXiv · Computation and Language · 13 Aug 2026 · paper
Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. Howev…
arXiv · Computation and Language · 13 Aug 2026 · paper
Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers…
arXiv · Machine Learning · 13 Aug 2026 · paper
Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using G…
arXiv · Machine Learning · 13 Aug 2026 · paper
Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore, recent research has increasing…
arXiv · Machine Learning · 13 Aug 2026 · paper
Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over re…
arXiv · Computation and Language · 13 Aug 2026 · paper
Proposal. Long context can replay history, but it does not decide which completed observations deserve authority. MMLA formalizes a bounded resident memory between transient context and slo…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language models (LLMs) show remarkable flexibility in adapting to novel tasks without parameter updates, a capacity known as in-context learning (ICL). Prior work has sought to unders…
arXiv · Computation and Language · 13 Aug 2026 · paper
Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing s…
arXiv · Computation and Language · 13 Aug 2026 · paper
Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. We instead route using internal L…
arXiv · Machine Learning · 13 Aug 2026 · paper
Low-precision formats usually optimize scalar fidelity while inheriting conventional product arithmetic. We introduce CurveFP, a block-scaled family that distributes magnitudes across inter…
arXiv · Machine Learning · 13 Aug 2026 · paper
Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by generic statistical priors. Richer do…
arXiv · Machine Learning · 13 Aug 2026 · paper
Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a surrogate loss as a pro…
arXiv · Machine Learning · 13 Aug 2026 · paper
Learning algorithms can be significantly improved by routing complex or uncertain inputs to specialized experts, balancing accuracy with computational cost. This approach, known as learning…
arXiv · Machine Learning · 13 Aug 2026 · paper
Training large-scale generative models is resource-intensive and relies heavily on heuristic dataset weighting. We address two fundamental questions: Can we train Large Language Models (LLM…
arXiv · Machine Learning · 13 Aug 2026 · paper
Designing generalizable control policies that operate reliably under changing conditions is essential for robust network services in modern digital infrastructure. Yet network control remai…
arXiv · Machine Learning · 13 Aug 2026 · paper
Time series forecasting drives operational decisions under tight latency budgets, and autoregressive time series foundation models (TSFMs) increasingly deliver the most accurate forecasts.…
arXiv · Machine Learning · 13 Aug 2026 · paper
Low-Rank Adaptation (LoRA) has become a popular technique for parameter-efficient fine-tuning of large language models (LLMs). In many real-world scenarios, multiple adapters are loaded sim…
arXiv · Machine Learning · 13 Aug 2026 · paper
Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that contin…
arXiv · Machine Learning · 13 Aug 2026 · paper
Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. However, re…
arXiv · Computation and Language · 13 Aug 2026 · paper
Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM…
arXiv · Computation and Language · 13 Aug 2026 · paper
A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a…
arXiv · Computation and Language · 13 Aug 2026 · paper
Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. We explore the cost effective adaptation of agents to novel domains by means of learning skills.…
arXiv · Machine Learning · 13 Aug 2026 · paper
A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction c…
arXiv · Computation and Language · 13 Aug 2026 · paper
We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains…
arXiv · Machine Learning · 13 Aug 2026 · paper
This paper reformulates Transformer/Attention mechanisms in Large Language Models (LLMs) through measure theory and frequency analysis, theoretically demonstrating that hallucination is an…
arXiv · Machine Learning · 13 Aug 2026 · paper
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence dat…
arXiv · Machine Learning · 13 Aug 2026 · paper
Post-training quantization pipelines routinely leave the softmax output layer in high precision. Yet in small LLMs with modern vocabularies, the head holds 15--30\% of all parameters, so a…
arXiv · Computation and Language · 13 Aug 2026 · paper
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledge from a…
arXiv · Machine Learning · 13 Aug 2026 · paper
Model merging by task arithmetic works until it doesn't, and the field diagnoses why with magnitudes: layerwise representation bias, deviations from cross-task linearity, parameter overlap.…
arXiv · Machine Learning · 13 Aug 2026 · paper
Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary acr…
arXiv · Machine Learning · 13 Aug 2026 · paper
Recent advances in EEG foundation models have demonstrated the potential of large-scale pretraining to enable generalizable neural decoding across subjects, recording environments, and data…