Daily Arxiv

전 세계에서 발간되는 인공지능 관련 논문을 정리하는 페이지 입니다.
본 페이지는 Google Gemini를 활용해 요약 정리하며, 비영리로 운영 됩니다.
논문에 대한 저작권은 저자 및 해당 기관에 있으며, 공유 시 출처만 명기하면 됩니다.

Rich Vehicle Routing Problem in Disaster Management enabling Temporally-causal Transhipments across Multi-Modal Transportation Network

Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering

Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations

The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models

Standards in the Preparation of Biomedical Research Metadata: A Bridge2AI Perspective

Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems

Zero-Knowledge Proofs in Sublinear Space

Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem

Training Text-to-Molecule Models with Context-Aware Tokenization

Towards Trustworthy Vital Sign Forecasting: Leveraging Uncertainty for Prediction Intervals

Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation

Language Models Identify Ambiguities and Exploit Loopholes

Mining the Long Tail: A Comparative Study of Data-Centric Criticality Metrics for Robust Offline Reinforcement Learning in Autonomous Motion Planning

GWM: Towards Scalable Gaussian World Models for Robotic Manipulation

ECHO: Frequency-aware Hierarchical Encoding for Variable-length Signals

Hierarchical Evaluation Function: A Multi-Metric Approach for Optimizing Demand Forecasting Models

Posterior-GRPO: Rewarding Reasoning Processes in Code Generation

Legal Knowledge Graph Foundations, Part I: URI-Addressable Abstract Works (LRMoo F1 to schema.org)

Pareto-Grid-Guided Large Language Models for Fast and High-Quality Heuristics Design in Multi-Objective Combinatorial Optimization

Self-supervised learning on gene expression data

EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model

Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques

MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform

Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies

Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint

Scaling Up Liquid-Resistance Liquid-Capacitance Networks for Efficient Sequence Modeling

From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

Defending against Indirect Prompt Injection by Instruction Detection

Using LLMs in Generating Design Rationale for Software Architecture Decisions

Evolution Meets Diffusion: Efficient Neural Architecture Generation

Direct Video-Based Spatiotemporal Deep Learning for Cattle Lameness Detection

FedDiverse: Tackling Data Heterogeneity in Federated Learning with Diversity-Driven Client Selection

Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation

Catch Me if You Search: When Contextual Web Search Results Affect the Detection of Hallucinations

COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing

ComfyGPT: A Self-Optimizing Multi-Agent System for Comprehensive ComfyUI Workflow Generation

CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning

CoPL: Collaborative Preference Learning for Personalizing LLMs

Structured Preference Optimization for Vision-Language Long-Horizon Task Planning

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

LocalEscaper: A Weakly-supervised Framework with Regional Reconstruction for Scalable Neural TSP Solvers

Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon

Gemstones: A Model Suite for Multi-Faceted Scaling Laws

A Deep Learning Pipeline for Solid Waste Detection in Remote Sensing Images

Learning Temporal Invariance in Android Malware Detectors

Beyond checkmate: exploring the creative chokepoints in AI text

Enhancing the De-identification of Personally Identifiable Information in Educational Data

LLM-ABBA: Understanding time series via symbolic approximation

Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland

Mirror-Consistency: Harnessing Inconsistency in Majority Voting

DAVIS: Planning Agent with Knowledge Graph-Powered Inner Monologue

Semantic Alignment-Enhanced Code Translation via an LLM-Based Multi-Agent System

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

AppAgent v2: Advanced Agent for Flexible Mobile Interactions

xGen-MM (BLIP-3): A Family of Open Large Multimodal Models

Towards Unified and Adaptive Cross-Domain Collaborative Filtering via Graph Signal Processing

Self-adaptive weights based on balanced residual decay rate for physics-informed neural networks and deep operator networks

Database-Augmented Query Representation for Information Retrieval

Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics

Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts

Empowering Time Series Analysis with Foundation Models: A Comprehensive Survey

FedCoSR: Personalized Federated Learning with Contrastive Shareable Representations for Label Heterogeneity in Non-IID Data

Conformal Temporal Logic Planning using Large Language Models

Learn to Relax with Large Language Models: Solving Nonlinear Combinatorial Optimization Problems via Bidirectional Coevolution

Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives

Caught in the Act: a mechanistic approach to detecting deception

When Truthful Representations Flip Under Deceptive Instructions?

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary

MAFA: A multi-agent framework for annotation

Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation

Designing AI-Agents with Personalities: A Psychometric Approach

Apertus: Democratizing Open and Compliant LLMs for Global Language Environments

Language models' activations linearly encode training-order recency

A Universal Banach--Bregman Framework for Stochastic Iterations: Unifying Stochastic Mirror Descent, Learning and LLM Training

Dense Video Understanding with Gated Residual Tokenization

Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting

Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs

TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning

Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions

Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework

Queen Detection in Beehives via Environmental Sensor Fusion for Low-Power Edge Computing

Machines are more productive than humans until they aren't, and vice versa

Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices

Prompt2Auto: From Motion Prompt to Automated Control via Geometry-Invariant One-Shot Gaussian Process Learning

PhenoGnet: A Graph-Based Contrastive Learning Framework for Disease Similarity Prediction

SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation

You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models

Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency

Differential Privacy in Federated Learning: Mitigating Inference Attacks with Randomized Response

LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology

An Empirical Study on Failures in Automated Issue Solving

DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models

MAP: End-to-End Autonomous Driving with Map-Assisted Planning

Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework

Created by

Haebom

저자

Kerui Huang, Shuhan Liu, Xing Hu, Tongtong Xu, Lingfeng Bao, Xin Xia

개요

본 논문은 사고 과정(Chain-of-Thought, CoT) 추론을 사용하여 대규모 언어 모델(LLM)의 성능을 향상시키는 연구에 대해 다룹니다. CoT는 중간 단계를 제시하여 산술, 논리, 상식적 과제에서 정확성과 견고성을 향상시키지만, 높은 계산 비용이라는 단점이 있습니다. 특히 간결하고 결정적인 출력이 필요한 소프트웨어 엔지니어링 작업에서는 이 문제가 더욱 심각합니다. 본 연구는 코드 생성 벤치마크를 기반으로 실험적 연구를 수행하여 과도한 CoT 추론이 정확도 저하, 지연 시간 증가, 출력 잘림 등의 문제를 야기함을 밝혔습니다. 이를 해결하기 위해, 본 논문은 정확도를 유지하면서 CoT를 압축하는 적응형 프레임워크인 SEER(Self-Enhancing Efficient Reasoning)을 제안합니다. SEER은 Best-of-N 샘플링과 작업별 적응형 필터링을 결합하여 사전 추론 출력에 따라 역치를 동적으로 조정하여 장황함과 계산 오버헤드를 줄입니다. 소프트웨어 엔지니어링 작업 세 가지와 수학 작업 한 가지에 대한 평가 결과, SEER은 CoT를 평균 42.1% 단축하고, 정확도를 향상시키며, 무한 루프를 대부분 제거하는 것으로 나타났습니다.

시사점, 한계점

•

시사점:

◦

CoT 추론의 효율성을 향상시키는 SEER 프레임워크 제시.

◦

과도한 CoT 추론의 부정적 영향(정확도 저하, 지연 시간 증가, 출력 잘림)을 실험적으로 입증.

◦

적응형 CoT 제어의 필요성 강조.

◦

SEER을 통해 CoT 기반 LLM의 효율성 및 견고성 향상 가능성 제시.

•

한계점:

◦

SEER의 성능은 사용되는 LLM과 작업의 특성에 따라 달라질 수 있음.

◦

제한된 벤치마크 데이터셋을 사용하여 일반화 성능에 대한 추가 연구 필요.

◦

다양한 유형의 소프트웨어 엔지니어링 작업에 대한 더욱 광범위한 평가 필요.

Made with Slashpage