Skip to main content
Scour
Discover
Docs
Login
Sign Up
You are offline. Trying to reconnect...
Copied to clipboard
Unable to share or copy to clipboard
cs updates on arXiv.org
rss.arxiv.org
arXiv
·
9h
9 hours ago
Dynamically Allocating Evaluation Effort for Model Ranking
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Dynamically Allocating Evaluation Effort for Model Ranking
arXiv
·
9h
9 hours ago
What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents
arXiv
·
9h
9 hours ago
Mapping the City Through the Lens of Language Models
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Mapping the City Through the Lens of Language Models
arXiv
·
9h
9 hours ago
Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning
arXiv
·
9h
9 hours ago
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
arXiv
·
9h
9 hours ago
Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks
arXiv
·
9h
9 hours ago
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
arXiv
·
9h
9 hours ago
GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
arXiv
·
9h
9 hours ago
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
arXiv
·
9h
9 hours ago
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
arXiv
·
9h
9 hours ago
BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems
arXiv
·
9h
9 hours ago
ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference
arXiv
·
9h
9 hours ago
Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
Covers
Benchmark & Compare the Best AI Models
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
arXiv
·
9h
9 hours ago
Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
arXiv
·
9h
9 hours ago
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
arXiv
·
9h
9 hours ago
HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification
arXiv
·
9h
9 hours ago
Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning
arXiv
·
9h
9 hours ago
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
arXiv
·
9h
9 hours ago
SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
arXiv
·
9h
9 hours ago
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
Love
Like
Not for me
Save
See related topics
Feeds
Share
Report
Spam
Misleading
Harmful Content
Block Domain
Actions for Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
« Page 14
·
Page 16 »
Log in to enable infinite scrolling
Keyboard Shortcuts
Navigation
Next / previous post
j
/
k
Open post
o
or
Enter
Preview post
v
Post Actions
Love post
a
Like post
l
Dislike post
d
Undo reaction
u
Save / unsave
s
Recommendations
Add interest / feed
Enter
Not interested
x
Go to
Home
g
h
Interests
g
i
Feeds
g
f
Likes
g
l
History
g
y
Changelog
g
c
Settings
g
s
Discover
g
b
Search
/
Pagination
Next page
n
Previous page
p
General
Show this help
?
Submit feedback
!
Close modal / unfocus
Esc
Press
?
anytime to show this help
Like
Save
Not for me
Report