Daily Arxiv

전 세계에서 발간되는 인공지능 관련 논문을 정리하는 페이지 입니다.
본 페이지는 Google Gemini를 활용해 요약 정리하며, 비영리로 운영 됩니다.
논문에 대한 저작권은 저자 및 해당 기관에 있으며, 공유 시 출처만 명기하면 됩니다.

Rich Vehicle Routing Problem in Disaster Management enabling Temporally-causal Transhipments across Multi-Modal Transportation Network

Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering

Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations

The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models

Standards in the Preparation of Biomedical Research Metadata: A Bridge2AI Perspective

Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems

Zero-Knowledge Proofs in Sublinear Space

Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem

Training Text-to-Molecule Models with Context-Aware Tokenization

Towards Trustworthy Vital Sign Forecasting: Leveraging Uncertainty for Prediction Intervals

Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation

Language Models Identify Ambiguities and Exploit Loopholes

Mining the Long Tail: A Comparative Study of Data-Centric Criticality Metrics for Robust Offline Reinforcement Learning in Autonomous Motion Planning

GWM: Towards Scalable Gaussian World Models for Robotic Manipulation

ECHO: Frequency-aware Hierarchical Encoding for Variable-length Signals

Hierarchical Evaluation Function: A Multi-Metric Approach for Optimizing Demand Forecasting Models

Posterior-GRPO: Rewarding Reasoning Processes in Code Generation

Legal Knowledge Graph Foundations, Part I: URI-Addressable Abstract Works (LRMoo F1 to schema.org)

Pareto-Grid-Guided Large Language Models for Fast and High-Quality Heuristics Design in Multi-Objective Combinatorial Optimization

Self-supervised learning on gene expression data

EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model

Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques

MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform

Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies

Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint

Scaling Up Liquid-Resistance Liquid-Capacitance Networks for Efficient Sequence Modeling

From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

Defending against Indirect Prompt Injection by Instruction Detection

Using LLMs in Generating Design Rationale for Software Architecture Decisions

Evolution Meets Diffusion: Efficient Neural Architecture Generation

Direct Video-Based Spatiotemporal Deep Learning for Cattle Lameness Detection

FedDiverse: Tackling Data Heterogeneity in Federated Learning with Diversity-Driven Client Selection

Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation

Catch Me if You Search: When Contextual Web Search Results Affect the Detection of Hallucinations

COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing

ComfyGPT: A Self-Optimizing Multi-Agent System for Comprehensive ComfyUI Workflow Generation

CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning

CoPL: Collaborative Preference Learning for Personalizing LLMs

Structured Preference Optimization for Vision-Language Long-Horizon Task Planning

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

LocalEscaper: A Weakly-supervised Framework with Regional Reconstruction for Scalable Neural TSP Solvers

Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon

Gemstones: A Model Suite for Multi-Faceted Scaling Laws

A Deep Learning Pipeline for Solid Waste Detection in Remote Sensing Images

Learning Temporal Invariance in Android Malware Detectors

Beyond checkmate: exploring the creative chokepoints in AI text

Enhancing the De-identification of Personally Identifiable Information in Educational Data

LLM-ABBA: Understanding time series via symbolic approximation

Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland

Mirror-Consistency: Harnessing Inconsistency in Majority Voting

DAVIS: Planning Agent with Knowledge Graph-Powered Inner Monologue

Semantic Alignment-Enhanced Code Translation via an LLM-Based Multi-Agent System

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

AppAgent v2: Advanced Agent for Flexible Mobile Interactions

xGen-MM (BLIP-3): A Family of Open Large Multimodal Models

Towards Unified and Adaptive Cross-Domain Collaborative Filtering via Graph Signal Processing

Self-adaptive weights based on balanced residual decay rate for physics-informed neural networks and deep operator networks

Database-Augmented Query Representation for Information Retrieval

Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics

Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts

Empowering Time Series Analysis with Foundation Models: A Comprehensive Survey

FedCoSR: Personalized Federated Learning with Contrastive Shareable Representations for Label Heterogeneity in Non-IID Data

Conformal Temporal Logic Planning using Large Language Models

Learn to Relax with Large Language Models: Solving Nonlinear Combinatorial Optimization Problems via Bidirectional Coevolution

Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives

Caught in the Act: a mechanistic approach to detecting deception

When Truthful Representations Flip Under Deceptive Instructions?

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary

MAFA: A multi-agent framework for annotation

Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation

Designing AI-Agents with Personalities: A Psychometric Approach

Apertus: Democratizing Open and Compliant LLMs for Global Language Environments

Language models' activations linearly encode training-order recency

A Universal Banach--Bregman Framework for Stochastic Iterations: Unifying Stochastic Mirror Descent, Learning and LLM Training

Dense Video Understanding with Gated Residual Tokenization

Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting

Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs

TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning

Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions

Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework

Queen Detection in Beehives via Environmental Sensor Fusion for Low-Power Edge Computing

Machines are more productive than humans until they aren't, and vice versa

Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices

Prompt2Auto: From Motion Prompt to Automated Control via Geometry-Invariant One-Shot Gaussian Process Learning

PhenoGnet: A Graph-Based Contrastive Learning Framework for Disease Similarity Prediction

SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation

You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models

Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency

Differential Privacy in Federated Learning: Mitigating Inference Attacks with Randomized Response

LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology

An Empirical Study on Failures in Automated Issue Solving

DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models

MAP: End-to-End Autonomous Driving with Map-Assisted Planning

Posterior-GRPO: Rewarding Reasoning Processes in Code Generation

Created by

Haebom

저자

Lishui Fan, Yu Zhang, Mouxiang Chen, Zhongxin Liu

개요

본 논문은 강화 학습(RL)을 이용한 대규모 언어 모델(LLM)의 코드 생성에서 중간 추론 과정의 질을 고려하는 새로운 프레임워크를 제시한다. 기존의 결과 기반 보상 방식의 한계를 극복하기 위해, 추론 과정의 질을 평가하는 벤치마크 LCB-RB와 추론 품질을 정확하게 평가하는 OD-based 방법을 제안한다. OD-based 방법은 추론 경로를 체계적으로 최적화 및 저하시켜 고품질 선호도 쌍을 생성한다. 또한, 성공적인 결과의 추론 과정에만 보상을 적용하는 새로운 RL 방법인 Posterior-GRPO(P-GRPO)를 제안하여 보상 해킹 문제를 완화한다. 7B 파라미터 모델을 사용한 실험 결과, P-GRPO는 다양한 코드 생성 작업에서 기존 방법보다 4.5% 향상된 성능을 보이며, GPT-4-Turbo와 비슷한 성능을 달성했다. 수학적 문제에도 적용 가능성을 보였다. 모델, 데이터셋, 코드는 공개적으로 이용 가능하다.

시사점, 한계점

•

시사점:

◦

중간 추론 과정의 질을 고려하는 새로운 강화 학습 프레임워크를 제시하여 LLM 기반 코드 생성 성능을 향상시켰다.

◦

추론 과정 평가를 위한 새로운 벤치마크 LCB-RB와 보상 모델 학습 방법 OD-based를 제안하여 추론 품질 평가의 정확성을 높였다.

◦

보상 해킹 문제를 완화하는 새로운 RL 알고리즘 P-GRPO를 제안했다.

◦

제안된 방법이 다양한 코드 생성 작업 및 수학적 문제에 일반화될 수 있음을 보였다.

◦

모델, 데이터셋, 코드를 공개하여 연구의 재현성과 확장성을 높였다.

•

한계점:

◦

LCB-RB 벤치마크의 범용성 및 확장성에 대한 추가적인 연구가 필요하다.

◦

OD-based 방법의 최적화 및 저하 전략의 개선 여지가 있다.

◦

P-GRPO 알고리즘의 복잡성 및 계산 비용에 대한 고려가 필요하다.

◦

현재 성능은 GPT-4-Turbo와 유사하지만, GPT-4-Turbo를 명확히 능가한다고 주장하기에는 추가적인 실험이 필요할 수 있다.

Made with Slashpage