arXiv:2605.14164v1 Announce Type: new Abstract: The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases…
cyberintel.kalymoon.com · 39366 articles · updated every 4 hours · grows forever
arXiv:2605.14164v1 Announce Type: new Abstract: The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases…
arXiv:2605.14163v1 Announce Type: new Abstract: Can a committee of weak reasoning-model calls reach the performance of much stronger models? We study verifier-backed committee search as inference-time…
arXiv:2605.14141v1 Announce Type: new Abstract: We study learning when the learned object is executable solver code rather than a predictor. In this setting, correctness is not enough: two solvers may…
arXiv:2605.14133v1 Announce Type: new Abstract: Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend…
arXiv:2605.14111v1 Announce Type: new Abstract: Hospital pharmacists make high-stakes decisions to mitigate drug shortages under uncertainty, time pressure, and patient risk. Interviews revealed that …
arXiv:2605.14102v1 Announce Type: new Abstract: Autonomous language-model agents increasingly combine planning, tool use, document processing, browsing, code execution, and verification loops. These c…
arXiv:2605.14089v1 Announce Type: new Abstract: In recent years, a variety of powerful LLM-based agentic systems have been applied to automate complex tasks through task orchestration. However, existi…
arXiv:2605.14062v1 Announce Type: new Abstract: While synthetic data generation with large language models (LLMs) is widely used in post-training pipelines, existing approaches typically generate full…
arXiv:2605.14061v1 Announce Type: new Abstract: Current autoformalization benchmarks are largely focused on olympiad or undergraduate mathematics, while graduate and research-level mathematics remains…
arXiv:2605.14054v1 Announce Type: new Abstract: Achieving robust perception-reasoning synergy is a central goal for advanced Vision-Language Models (VLMs). Recent advancements have pursued this goal v…
arXiv:2605.14051v1 Announce Type: new Abstract: Industrial LLM agent systems often separate planning from execution, yet LLM planners frequently produce structurally invalid or unnecessarily long work…
arXiv:2605.14049v1 Announce Type: new Abstract: The growing adoption of large language models in legal practice brings both significant promise and serious risk. Legal professionals stand to benefit f…
arXiv:2605.14048v1 Announce Type: new Abstract: Masked autoencoders (MAEs) have recently shown promise for self-supervised representation learning of resting-state brain functional connectivity (FC). …
arXiv:2605.14038v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as autonomous agents that must decide when to answer directly vs. when to invoke external tools. Prior wor…
arXiv:2605.14036v1 Announce Type: new Abstract: In current Large Language Models we can trust the production of smoothly flowing prose on the basis of the principles of machine learning. However, ther…
arXiv:2605.14034v1 Announce Type: new Abstract: Wide applications of LLM-based agents require strong alignment with human social values. However, current works still exhibit deficiencies in self-cogni…
arXiv:2605.14033v1 Announce Type: new Abstract: Scientific theory shift in AI agents requires more than fitting equations to data. An artificial scientific agent must detect whether an existing repres…
arXiv:2605.14004v1 Announce Type: new Abstract: Generative models are often trained with a next-token prediction objective, yet many downstream applications require the ability to estimate or control …
A vulnerability was found in AMD Radeon RX 6000 Graphics Products, Radeon PRO W6000 Graphics Products, Radeon PRO V520 and Radeon PRO V620 and classified as problematic . Impacted is an unknown functi…
A vulnerability was found in AMD Radeon PRO V710 . It has been classified as critical . The affected element is an unknown function. This manipulation causes improper isolation of shared resources on …
A vulnerability was found in AMD Ryzen 7035 Processors with Radeon Graphics, Ryzen 7040 Mobile Processors with Radeon Graphics, Ryzen 8040 Mobile Processors with Radeon Graphics, Ryzen 6000 Processors…
A vulnerability was found in AMD Ryzen 7040 Mobile Processors with Radeon Graphics, Ryzen 8000 Desktop Processors, Ryzen 8040 Mobile Processors with Radeon Graphics and Ryzen Embedded 8000 Processors …
A vulnerability categorized as problematic has been discovered in AMD Ryzen 5000 Desktop Processors with Radeon Graphics, Ryzen Threadripper PRO 5000 WX-Series Processors, Ryzen 7030 Mobile Processors…
A vulnerability identified as problematic has been detected in AMD Ryzen Al Max+, Ryzen AI 300 Processors, Ryzen 7040 Mobile Processors with Radeon Graphics, Ryzen 8000 Desktop Processors, Ryzen 8040 …