Daily Arxiv

전 세계에서 발간되는 인공지능 관련 논문을 정리하는 페이지 입니다.
본 페이지는 Google Gemini를 활용해 요약 정리하며, 비영리로 운영 됩니다.
논문에 대한 저작권은 저자 및 해당 기관에 있으며, 공유 시 출처만 명기하면 됩니다.

Arbitrary Precision Printed Ternary Neural Networks with Holistic Evolutionary Approximation

Invited Paper: Feature-to-Classifier Co-Design for Mixed-Signal Smart Flexible Wearables for Healthcare at the Extreme Edge

Robustness is Important: Limitations of LLMs for Data Fitting

CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens

CE-RS-SBCIT A Novel Channel Enhanced Hybrid CNN Transformer with Residual, Spatial, and Boundary-Aware Learning for Brain Tumor MRI Analysis

PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science

THEME: Enhancing Thematic Investing with Semantic Stock Representations and Temporal Dynamics

Trust but Verify! A Survey on Verification Design for Test-time Scaling

Quantized Neural Networks for Microcontrollers: A Comprehensive Review of Methods, Platforms, and Applications

Documenting Deployment with Fabric: A Repository of Real-World AI Governance

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

Region-Level Context-Aware Multimodal Understanding

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism

Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention

Adaptive Duration Model for Text Speech Alignment

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs

Time-RA: Towards Time Series Reasoning for Anomaly with LLM Feedback

Dually Hierarchical Drift Adaptation for Online Configuration Performance Learning

Single Domain Generalization for Multimodal Cross-Cancer Prognosis via Dirac Rebalancer and Distribution Entanglement

Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization

Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models

Scientifically-Interpretable Reasoning Network (ScIReN): Discovering Hidden Relationships in the Carbon Cycle and Beyond

A Hybrid Artificial Intelligence Method for Estimating Flicker in Power Systems

Beyond Frequency: The Role of Redundancy in Large Language Model Memorization

TrueGL: A Truthful, Reliable, and Unified Engine for Grounded Learning in Full-Stack Search

Unified Path Planner with Adaptive Safety and Optimality

FedSEA-LLaMA: A Secure, Efficient and Adaptive Federated Splitting Framework for Large Language Models

WebInject: Prompt Injection Attack to Web Agents

Towards Embodiment Scaling Laws in Robot Locomotion

SPIN-ODE: Stiff Physics-Informed Neural ODE for Chemical Reaction Rate Estimation

DDaTR: Dynamic Difference-aware Temporal Residual Network for Longitudinal Radiology Report Generation

Latent Adaptive Planner for Dynamic Manipulation

MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awareness

SAGA: A Security Architecture for Governing AI Agentic Systems

Towards Understanding Camera Motions in Any Video

Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction

DeepTrans: Deep Reasoning Translation via Reinforcement Learning

A Hybrid Fully Convolutional CNN-Transformer Model for Inherently Interpretable Disease Detection from Retinal Fundus Images

Decentralized Domain Generalization with Style Sharing: Formal Model and Convergence Analysis

FROG: Fair Removal on Graphs

DPImageBench: A Unified Benchmark for Differentially Private Image Synthesis

LLM Test Generation via Iterative Hybrid Program Analysis

Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts

Retrieval-Augmented Machine Translation with Unstructured Knowledge

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning

RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis

Categorical Data Clustering via Value Order Estimated Distance Metric Learning

Guiding a diffusion model using sliding windows

A Collaborative Content Moderation Framework for Toxicity Detection based on Conformalized Estimates of Annotation Disagreement

Mamba State-Space Models Are Lyapunov-Stable Learners

Alice's Adventures in a Differentiable Wonderland -- Volume I, A Tour of the Land

COBRA-PPM: A Causal Bayesian Reasoning Architecture Using Probabilistic Programming for Robot Manipulation Under Uncertainty

Large Intestine 3D Shape Refinement Using Point Diffusion Models for Digital Phantom Generation

What Breaks Knowledge Graph based RAG? Empirical Insights into Reasoning under Incomplete Knowledge

QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges

AI Simulation by Digital Twins: Systematic Survey, Reference Framework, and Mapping to a Standardized Architecture

Compression versus Accuracy: A Hierarchy of Lifted Models

TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving

Evaluating Knowledge Graph Based Retrieval Augmented Generation Methods under Knowledge Incompleteness

Transforming Wearable Data into Personal Health Insights using Large Language Model Agents

Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning

DynaMark: A Reinforcement Learning Framework for Dynamic Watermarking in Industrial Machine Tool Controllers

TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank

MoE-Health: A Mixture of Experts Framework for Robust Multimodal Healthcare Prediction

Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval

PiCSAR: Probabilistic Confidence Selection And Ranking

Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight

Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering

Reasoning-Intensive Regression

Neural Network Acceleration on MPSoC board: Integrating SLAC's SNL, Rogue Software and Auto-SNL

Developer Insights into Designing AI-Based Computer Perception Tools

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models

OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization

Entropy-Based Non-Invasive Reliability Monitoring of Convolutional Neural Networks

Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR

Harnessing IoT and Generative AI for Weather-Adaptive Learning in Climate Resilience Education

QZhou-Embedding Technical Report

Physics-Informed Spectral Modeling for Hyperspectral Imaging

Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning

A Survey on Current Trends and Recent Advances in Text Anonymization

NSPDI-SNN: An efficient lightweight SNN based on nonlinear synaptic pruning and dendritic integration

Limitations of Physics-Informed Neural Networks: a Study on Smart Grid Surrogation

EZ-Sort: Efficient Pairwise Comparison via Zero-Shot CLIP-Based Pre-Ordering and Human-in-the-Loop Sorting

What Data is Really Necessary? A Feasibility Study of Inference Data Minimization for Recommender Systems

Complete Gaussian Splats from a Single Image with Denoising Diffusion Models

On the Hardness of Learning GNN-based SAT Solvers: The Role of Graph Ricci Curvature

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning

HSFN: Hierarchical Selection for Fake News Detection building Heterogeneous Ensemble

Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards

Controllable 3D Molecular Generation for Structure-Based Drug Design Through Bayesian Flow Networks and Gradient Integration

Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate Prediction

MedShift: Implicit Conditional Transport for X-Ray Domain Adaptation

The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models

Benchmarking the State of Networks with a Low-Cost Method Based on Reservoir Computing

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

Created by

Haebom

저자

Peng Xia, Jinglu Wang, Yibo Peng, Kaide Zeng, Xian Wu, Xiangru Tang, Hongtu Zhu, Yun Li, Shujie Liu, Yan Lu, Huaxiu Yao

개요

본 논문은 다양한 의료 전문 분야에 걸쳐 일반화하는 데 어려움을 겪는 기존의 단일 에이전트 의료 대규모 시각-언어 모델(Med-LVLMs)의 한계를 극복하기 위해, 강화 학습(RL) 기반 다중 에이전트 프레임워크인 MMedAgent-RL을 제안합니다. MMedAgent-RL은 환자를 적절한 전문 분야에 배정하는 분류 의사와 다중 전문가의 판단과 자체 지식을 통합하여 최종 결정을 내리는 주치의, 두 가지의 Qwen2.5-VL 기반 GP 에이전트로 구성됩니다. 전문가 출력의 불일치 문제 해결을 위해, 주치의가 전문가 모방과 실수 수정 간의 균형을 점진적으로 학습하도록 하는 커리큘럼 학습(CL) 기반 RL 전략을 도입했습니다. 다섯 가지 의료 VQA 벤치마크 실험 결과, MMedAgent-RL은 오픈소스 및 독점 Med-LVLMs를 능가하며, 사람과 유사한 추론 패턴을 보이는 것으로 나타났습니다. 특히, 지도 학습 기반 미세 조정 기준 모델 대비 평균 20.7%의 성능 향상을 달성했습니다.

시사점, 한계점

•

시사점:

◦

기존 단일 에이전트 Med-LVLMs의 한계를 극복하는 강화학습 기반 다중 에이전트 협업 프레임워크 제시

◦

동적이고 최적화된 다중 전문가 협업을 통한 의료 영상 분석 및 진단 성능 향상

◦

커리큘럼 학습을 통한 전문가 의견 불일치 문제 해결 및 인간 수준의 추론 패턴 구현

◦

기존 모델 대비 유의미한 성능 향상 (평균 20.7%) 달성

•

한계점:

◦

제안된 모델의 일반화 성능에 대한 추가적인 검증 필요

◦

다양한 의료 데이터셋에 대한 실험 결과 제시 부족

◦

실제 임상 환경 적용을 위한 추가적인 연구 필요

◦

Qwen2.5-VL 모델에 대한 의존성으로 인한 다른 언어 모델 적용의 어려움 또는 제약 존재 가능성

Made with Slashpage