Daily Arxiv

전 세계에서 발간되는 인공지능 관련 논문을 정리하는 페이지 입니다.
본 페이지는 Google Gemini를 활용해 요약 정리하며, 비영리로 운영 됩니다.
논문에 대한 저작권은 저자 및 해당 기관에 있으며, 공유 시 출처만 명기하면 됩니다.

ACCeLLiuM: Supervised Fine-Tuning for Automated OpenACC Pragma Generation

AnchDrive: Bootstrapping Diffusion Policies with Hybrid Trajectory Anchors for End-to-End Driving

Diffusion-Augmented Contrastive Learning: A Noise-Robust Encoder for Biosignal Representations

FusedANN: Convexified Hybrid ANN via Attribute-Vector Fusion

HiCoLoRA: Addressing Context-Prompt Misalignment via Hierarchical Collaborative LoRA for Zero-Shot DST

A Longitudinal Randomized Control Study of Companion Chatbot Use: Anthropomorphism and Its Mediating Role on Social Impacts

TimeMosaic: Temporal Heterogeneity Guided Time Series Forecasting via Adaptive Granularity Patch and Segment-wise Decoding

Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models

Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

Comparing RAG and GraphRAG for Page-Level Retrieval Question Answering on Math Textbook

RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation

Distribution-Aligned Decoding for Efficient LLM Task Adaptation

DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models

Recent Advancements in Microscopy Image Enhancement using Deep Learning: A Survey

Constructive Conflict-Driven Multi-Agent Reinforcement Learning for Strategic Diversity

Towards a Physics Foundation Model

Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews

Positional Encoding via Token-Aware Phase Attention

Chain or tree? Re-evaluating complex reasoning from the perspective of a matrix of thought

A Two-Stage Strategy for Mitosis Detection Using Improved YOLO11x Proposals and ConvNeXt Classification

JudgeAgent: Knowledge-wise and Dynamic LLM Evaluation with Agent-as-Interviewer

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

Scalable Option Learning in High-Throughput Environments

"She was useful, but a bit too optimistic": Augmenting Design with Interactive Virtual Personas

In-Context Algorithm Emulation in Fixed-Weight Transformers

Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

Conflict-Aware Soft Prompting for Retrieval-Augmented Generation

ERIS: An Energy-Guided Feature Disentanglement Framework for Out-of-Distribution Time Series Classification

StreetReaderAI: Making Street View Accessible Using Context-Aware Multimodal AI

Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning

Intuition emerges in Maximum Caliber models at criticality

GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy

Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models

Decentralized Aerial Manipulation of a Cable-Suspended Load using Multi-Agent Reinforcement Learning

SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy

DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS

Generative Logic: A New Computer Architecture for Deterministic Reasoning and Knowledge Generation

Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles

Hierarchical Graph Neural Network for Compressed Speech Steganalysis

R-Stitch: Dynamic Trajectory Stitching for Efficient Reasoning

The Invisible Leash: Why RLVR May or May Not Escape Its Origin

APTx Neuron: A Unified Trainable Neuron Architecture Integrating Activation and Computation

LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues

Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities

KV Cache Steering for Controlling Frozen LLMs

Lightweight MSA Design Advances Protein Folding From Evolutionary Embeddings

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders

Neural-Network solver of ideal MHD equilibria

Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective

Beyond Simple Graphs: Neural Multi-Objective Routing on Multigraphs

On the Necessity of Output Distribution Reweighting for Effective Class Unlearning

TAMMs: Temporal-Aware Multimodal Model for Satellite Image Change Understanding and Forecasting

Latent Concept Disentanglement in Transformer-based Language Models

Personalized LLM Decoding via Contrasting Personal Preference

Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training

Security Degradation in Iterative AI Code Generation -- A Systematic Analysis of the Paradox

Think With Videos For Agentic Long-Video Understanding

VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks

Position: Simulating Society Requires Simulating Thought

AMPED: Adaptive Multi-objective Projection for balancing Exploration and skill Diversification

DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models

Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement

Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

Physics-Guided Motion Loss for Video Generation Model

Probing Neural Topology of Large Language Models

InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning

Mamba Integrated with Physics Principles Masters Long-term Chaotic System Forecasting

Model-Preserving Adaptive Rounding

DORAEMON: Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation

SDPO: Importance-Sampled Direct Preference Optimization for Stable Diffusion Training

Spectral-inspired Operator Learning with Limited Data and Unknown Physics

BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases

Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training

Prompting is not Enough: Exploring Knowledge Integration and Controllable Generation on Large Language Models

HD-PiSSA: High-Rank Distributed Orthogonal Adaptation

Can LLMs Alleviate Catastrophic Forgetting in Graph Continual Learning? A Systematic Study

FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models

From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning

BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation

Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning

The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm

Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs

Learning Flexible Forward Trajectories for Masked Molecular Diffusion

Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems

Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

UniErase: Towards Balanced and Precise Unlearning in Language Models

Octic Vision Transformers: Quicker ViTs Through Equivariance

Intentional Gesture: Deliver Your Intentions with Gestures for Speech

UltraEdit: Training-, Subject-, and Memory-Free Lifelong Editing in Language Models

VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation

Learning Hierarchical Domain Models Through Environment-Grounded Interaction

Shadow-FT: Tuning Instruct Model via Training on Paired Base Model

Structured Relational Representations

Latent Veracity Inference for Identifying Errors in Stepwise Reasoning

Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders

ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training

VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks

Created by

Haebom

저자

Xinlong Chen, Yuanxing Zhang, Yushuo Guan, Weihong Lin, Zekun Wang, Bohan Zeng, Yang Shi, Sihan Yang, Qiang Liu, Pengfei Wan, Liang Wang, Tieniu Tan

개요

"Reason-Then-Respond" 패러다임을 강화 학습과 결합한 접근 방식은 Multimodal Large Language Models의 발전에 기여했으나, 비디오 도메인에 적용 시 질문 응답 (QA) 또는 캡셔닝 작업 중 하나에 특화된 모델을 양산하여 두 가지 작업을 모두 수행하는 데 어려움을 겪었다. 서로 상반된 작업 특성으로 인해 두 작업의 보상 신호를 단순히 결합하면 성능 저하가 발생한다. 이러한 문제를 해결하기 위해, 본 논문은 DarkEventInfer와 MixVidQA라는 두 가지 중간 프록시 작업을 기반으로 하는 새로운 학습 프레임워크를 제안한다. DarkEventInfer는 마스크 처리된 이벤트 세그먼트가 있는 비디오를 제시하여 모델이 컨텍스트 비디오 단서를 기반으로 가려진 내용을 추론하도록 요구하며, MixVidQA는 두 개의 다른 클립으로 구성된 인터리빙된 비디오 시퀀스를 제시하여 모델이 하나를 격리하고 추론하면서 다른 하나를 무시하도록 요구한다. 이 프레임워크를 통해 전체적이고 발산적인 이해와 정확하고 수렴적인 추론 능력을 동시에 개발하도록 유도한다. 이 프레임워크를 구현한 VidBridge-R1은 패러다임 충돌을 효과적으로 해결하는 최초의 다목적 비디오 추론 모델이다. 광범위한 실험을 통해 VidBridge-R1이 하나의 모델 내에서 QA 및 캡셔닝 모두에서 상당한 성능 향상을 달성했으며, 보다 일반화되고 강력한 비디오 이해 모델을 육성하는 데 있어 제안된 접근 방식의 효과를 입증했다.

시사점, 한계점

•

시사점:

◦

단일 모델 내에서 QA 및 캡셔닝 작업 모두에서 상당한 성능 향상을 달성했다.

◦

비디오 이해 모델의 일반화 및 성능 향상에 기여했다.

◦

패러다임 충돌 문제를 해결하는 새로운 학습 프레임워크를 제시했다.

◦

DarkEventInfer와 MixVidQA라는 새로운 프록시 작업을 통해 모델의 이해 능력을 향상시켰다.

•

한계점:

◦

논문 자체에서 명시된 한계점은 언급되지 않았다. (논문에 한계점 관련 내용이 없다는 뜻)

Made with Slashpage