/
/
Daily Arxiv
Daily Arxiv
世界中で発行される人工知能関連の論文をまとめるページです。
このページはGoogle Geminiを活用して要約し、非営利で運営しています。
論文の著作権は著者および関連機関にあり、共有する際は出典を明記してください。
Emotions as Ambiguity-aware Ordinal Representations
From Tabula Rasa to Emergent Abilities: Discovering Robot Skills via Real-World Unsupervised Quality-Diversity
Enhancing Model Privacy in Federated Learning with Random Masking and Quantization
Scaling Laws for Task-Stratified Knowledge in Post-Training Quantized Large Language Models
Principled Detection of Hallucinations in Large Language Models via Multiple Testing
Vocoder-Projected Feature Discriminator
ControlEchoSynth: Boosting Ejection Fraction Estimation Models via Controlled Video Diffusion
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Time-Aware One Step Diffusion Network for Real-World Image Super-Resolution
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives
GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
Input-Time Scaling
LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
StreetViewAI: Making Street View Accessible Using Context-Aware Multimodal AI
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
From Imitation to Optimization: A Comparative Study of Offline Learning for Autonomous Driving
R-Zero: Self-Evolving Reasoning LLM from Zero Data
Human-Centered Human-AI Interaction (HC-HAII): A Human-Centered AI パースペクティブ
GTPO: Trajectory-Based Policy Optimization in Large Language Models
Contrastive Multi-Task Learning with Solvent-Aware Augmentation for Drug Discovery
A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics
Invisible Architectures of Thought: Toward a New Science of AI as Cognitive Infrastructure
Revisiting Pre-trained Language Models for Vulnerability Detection
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Scaling Decentralized Learning with FLock
SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
Apple Intelligence Foundation Language Models: Tech Report 2025
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
PyVision: Agentic Vision with Dynamic Tooling
DATABench: Evaluating Dataset Auditing in Deep Learning from an Adversarial Perspective
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
Analyzing Character Representation in Media Content using Multimodal Foundation Model: Effectiveness and Trust
MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
Pseudo-Simulation for Autonomous Driving
BinConv: A Neural Architecture for Ordinal Encoding in Time-Series Forecasting
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
EnvInjection: Environmental Prompt Injection Attack to Multi-modal Web Agents
X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real
Heat Diffusion Models - Interpixel Attention Mechanism
Bidirectional Task-Motion Planning Based on Hierarchical Reinforcement Learning for Strategic Confrontation
Multi-Type Context-Aware Conversational Recommender Systems via Mixture-of-Experts
Pricing AI Model Accuracy
Evaluating the Fitness of Ontologies for the Task of Question Generation
Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
PGAD: Prototype-Guided Adaptive Distillation for Multi-Modal Learning in AD Diagnosis
Constructing a Norm for Children's Scientific Drawing: Distribution Features Based on Semantic Similarity of Large Language Models
An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model
Efficient PINNs via Multi-Head Unimodular Regularization of the Solutions Space
Statistical learning does not always entail knowledge
Score-based Generative Diffusion Models for Social Recommendations
PromptKeeper: Safeguarding System Prompts for LLMs
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language モデル
Leveraging Multi-facet Paths for Heterogeneous Graph Representation Learning
Training with Explanations Alone: A New Paradigm to Prevent Shortcut Learning
Generation of Geodesics with Actor-Critic Reinforcement Learning to Predict Midpoints
TabSketchFM: Sketch-based Tabular Representation Learning for Data Discovery over Data Lakes
HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models
StepWiser: Stepwise Generative Judges for Wiser Reasoning
AniME: Adaptive Multi-Agent Planning for Long Animation Generation
AppAgent-Pro: A Proactive GUI Agent System for Multidomain Information Integration and User Assistance
AI Chaperones Are (Really) All You Need to Prevent Parasocial Relationships with Chatbots
Nemori: Self-Organizing Agent Memory Inspired by Cognitive Science
General agents contain world models
Approximate Lifted Model Construction
Fitness Landscape of Large Language Model-Assisted Automated Algorithm Search
Synthesizing High-Quality Programming Tasks with LLM-based Expert and Student Agents
Preference Elicitation for Multi-objective Combinatorial Optimization with Active Learning and Maximum Likelihood Estimation
Reference-Aligned Retrieval-Augmented Question Answering over Heterogeneous Proprietary Documents
Demonstrating specification gaming in reasoning models
AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation
Think Smart、Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning
From Evidence to Decision: Exploring Evaluative AI
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning
Discrete-Guided Diffusion for Scalable and Safe Multi-Robot Motion Planning
Patch Progression Masked Autoencoder with Fusion CNN Network for Classifying Evolution Between Two Pairs of 2D OCT Slices
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
Large Language Models (LLMs) for Electronic Design Automation (EDA)
Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence
HPC Digital Twins for Evaluating Scheduling Policies, Incentive Structures and their Impact on Power and Cooling
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
Cross-Platform E-Commerce Product Categorization and Recategorization: A Multimodal Hierarchical Classification Approach
Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
MathBuddy: A Multimodal System for Affective Math Tutoring
Diffusion Language Models Know the Answer Before Decoding
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation
WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution
The Next Layer: Augmenting Foundation Models with Structure-Preserving and Attention-Guided Learning for Local Patches to Global Context Awareness in Computational Pathology
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
The Information Dynamics of Generative Diffusion
AI-Powered Detection of Inappropriate Language in Medical School Curricula
Generative AI for Testing of Autonomous Driving Systems: A Survey
Multispectral LiDAR data for extracting tree points in urban and suburban areas
Load more
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning
Created by
Haebom
作者
Zeyi Sun, Yuhang Cao, Jianze Liang, Qiushi Sun, Ziyu Liu, Zhixiong Zhang, Yuhang Zang, Xiaoyi Dong, Kai Chen, Dahua Lin, Jiaqi Wang
概要
本論文は、科学コンピューティングなどの専門分野におけるグラフィカルユーザーインターフェース(GUI)のための自律エージェントの設計上の問題を解決するために、長期計画と正確な実行の両方が必要な状況で、既存の一般的なエージェントと専門的なエージェントの限界を克服する新しいアプローチを提示します。既存のアプローチは計画能力と実行能力との間に矛盾があるが、本論文で提示されているCODAは一般的な計画者(Cerebrum)と専門家の実行者(Cerebellum)を統合する学習可能な構成型フレームワークです。 CODAは2段階のパイプラインで訓練されます。最初のステップであるSpecializationでは、各科学アプリケーションに対して専門の計画者を個別にトレーニングし、2番目のステップであるGeneralizationはすべての成功した軌跡を集めて、最終計画者のための指導学習の微調整に使用します。これにより、CODAは強力な実行能力とドメイン間の一般化能力の両方を備えています。 ScienceBoardベンチマークの4つの課題では、CODAは既存の方法を大幅に上回り、オープンソースモデルの中で最高のパフォーマンスを達成します。
Takeaways、Limitations
•
Takeaways:
◦
科学コンピューティングの分野におけるGUI自律エージェントの性能向上のための新しいアプローチの提示
◦
一般的な計画能力と専門的な実行能力を組み合わせることで、既存の限界を克服
◦
学習可能な構成型フレームワークを通じて経験から適応可能
◦
限られたデータ環境でも効果的なパフォーマンスを実現
◦
オープンソースモデル中の最高性能記録
•
Limitations:
◦
提示されたフレームワークの一般化能力の追加評価が必要
◦
さまざまな科学分野とより複雑なGUI環境へのスケーラビリティ検証が必要
◦
ScienceBoardベンチマーク以外のベンチマークでのパフォーマンス評価が必要
◦
トレーニングデータの品質への依存度評価が必要
PDFを見る
Made with Slashpage