[공지사항]을 빙자한 안부와 근황

Show more

Daily Arxiv

전 세계에서 발간되는 인공지능 관련 논문을 정리하는 페이지 입니다.
본 페이지는 Google Gemini를 활용해 요약 정리하며, 비영리로 운영 됩니다.
논문에 대한 저작권은 저자 및 해당 기관에 있으며, 공유 시 출처만 명기하면 됩니다.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models

MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks

Site-Level Fine-Tuning with Progressive Layer Freezing: Towards Robust Prediction of Bronchopulmonary Dysplasia from Day-1 Chest Radiographs in Extremely Preterm Infants

A Roadmap for Climate-Relevant Robotics Research

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

MMOne: Representing Multiple Modalities in One Scene

SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance

(Almost) Free Modality Stitching of Foundation Models

A Brain Tumor Segmentation Method Based on CLIP and 3D U-Net with Cross-Modal Semantic Guidance and Multi-Level Feature Fusion

KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News Detection

THOR: Transformer Heuristics for On-Demand Retrieval

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems

KeyRe-ID: Keypoint-Guided Person Re-Identification using Part-Aware Representation in Videos

Prompt Perturbations Reveal Human-Like Biases in LLM Survey Responses

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Fast Bilateral Teleoperation and Imitation Learning Using Sensorless Force Control via Accurate Dynamics Model

Task-Specific Generative Dataset Distillation with Difficulty-Guided Sampling

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents

ReCode: Updating Code API Knowledge with Reinforcement Learning

Cross-Layer Discrete Concept Discovery for Interpreting Language Models

Semantic Structure-Aware Generative Attacks for Enhanced Adversarial Transferability

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents

Multiple-Frequencies Population-Based Training

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

GPU Performance Portability needs Autotuning

Generating Synthetic Data via Augmentations for Improved Facial Resemblance in DreamBooth and InstantID

Coral Protocol: Open Infrastructure Connecting The Internet of Agents

MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awareness

Federated Learning: A Survey on Privacy-Preserving Collaborative Intelligence

ConTextual: Improving Clinical Text Summarization in LLMs with Context-preserving Token Filtering and Knowledge Graphs

Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression

JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model

K-P Quantum Neural Networks

VectorFit : Adaptive Singular & Bias Vector Fine-Tuning of Pre-trained Foundation Models

Data-Efficient Deep Operator Network for Unsteady Flow: A Multi-Fidelity Approach with Physics-Guided Subsampling

Learning Universal Human Mobility Patterns with a Foundation Model for Cross-domain Data Fusion

GeoFlow-SLAM: A Robust Tightly-Coupled RGBD-Inertial and Legged Odometry Fusion SLAM for Dynamic Legged Robotics

A Multi-Stage Framework with Taxonomy-Guided Reasoning for Occupation Classification Using Large Language Models

Multi-View Node Pruning for Accurate Graph Representation

V-Max: A Reinforcement Learning Framework for Autonomous Driving

Interpretable Transformation and Analysis of Timelines through Learning via Surprisability

AI Governance InternationaL Evaluation Index (AGILE Index) 2024

UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning

Improving Transformer World Models for Data-Efficient RL

LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential Recommendation

SIDDA: SInkhorn Dynamic Domain Adaptation for Image Classification with Equivariant Neural Networks

Determination of galaxy photometric redshifts using Conditional Generative Adversarial Networks (CGANs)

Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis

MRGen: Segmentation Data Engine for Underrepresented MRI Modalities

IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization

Out-of-Distribution Recovery with Object-Centric Keypoint Inverse Policy for Visuomotor Imitation Learning

Dataset resulting from the user study on comprehensibility of explainable AI algorithms

Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models

LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization

Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information

DeFine: Decision-Making with Analogical Reasoning over Factor Profiles

Benchmarking Sub-Genre Classification For Mainstage Dance Music

Risks of ignoring uncertainty propagation in AI-augmented security pipelines

MedPix 2.0: A Comprehensive Multimodal Biomedical Data set for Advanced AI Applications with Retrieval Augmented Generation and Knowledge Graphs

Leveraging Quantum Superposition to Infer the Dynamic Behavior of a Spatial-Temporal Neural Network Signaling Model

Bounding the Worst-class Error: A Boosting Approach

TBDetector:Transformer-Based Detector for Advanced Persistent Threats with Provenance Graph

Machine Learning Systems: A Survey from a Data-Oriented Perspective

Aime: Towards Fully-Autonomous Multi-Agent Framework

SmartThinker: Learning to Compress and Preserve Reasoning by Step-Level Length Control

Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments

NTRL: Encounter Generation via Reinforcement Learning for Dynamic Difficulty Adjustment in Dungeons and Dragons

Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge

ActionStudio: A Lightweight Framework for Data and Training of Large Action Models

BEARCUBS: A benchmark for computer-using web agents

Demystifying MuZero Planning: Interpreting the Learned Model

LLM-Enhanced User-Item Interactions: Leveraging Edge Information for Optimized Recommendations

VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

Imbalance in Balance: Online Concept Balancing in Generation Models

Latent Policy Steering with Embodiment-Agnostic Pretrained World Models

Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It

Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark

AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research

Towards Formal Verification of LLM-Generated Code from Natural Language Prompts

Evaluating Reinforcement Learning Algorithms for Navigation in Simulated Robotic Quadrupeds: A Comparative Study Inspired by Guide Dog Behaviour

Overview of the TalentCLEF 2025: Skill and Job Title Intelligence for Human Capital Management

QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation

Merge Kernel for Bayesian Optimization on Permutation Space

Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy

Automating Steering for Safe Multimodal Large Language Models

HATS: Hindi Analogy Test Set for Evaluating Reasoning in Large Language Models

VITA: Vision-to-Action Flow Matching Policy

$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation

Synthesizing Reality: Leveraging the Generative AI-Powered Platform Midjourney for Construction Worker Detection

Aligning Humans and Robots via Reinforcement Learning from Implicit Human Feedback

SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks

Prompt Injection 2.0: Hybrid AI Threats

Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities

UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets

Created by

Haebom

저자

Zhichao Sheng, Shilin Zhou, Chen Gong, Zhenghua Li

개요

본 논문은 다양한 음성 중심 멀티미디어 애플리케이션에서 기계가 음성 언어를 이해할 수 있도록 하는 음성 언어 이해(SLU)에 초점을 맞추고 있습니다. SLU는 자동 음성 인식(ASR), 음성 개체명 인식(NER), 음성 감정 분석(SA) 등 여러 작업을 포함합니다. 기존 방법들은 각 작업에 대해 별도의 모델 아키텍처를 사용하여 시스템 복잡성을 증가시키고 작업 간 상호 작용을 제한하며 여러 작업에서 사용 가능한 이종 데이터 세트를 완전히 활용하지 못하는 한계를 가지고 있습니다. 본 논문에서는 이러한 한계를 해결하기 위해 단일 아키텍처 내에서 여러 SLU 작업을 공동으로 모델링하는 통합 프레임워크인 UniSLU를 제안합니다. UniSLU는 다양한 SLU 작업에 대한 통합된 표현을 제안하여 여러 작업에 걸쳐 이종 데이터 세트를 완전히 활용할 수 있도록 합니다. 이 표현을 기반으로 ASR, 음성 NER 및 SA 작업을 공동으로 모델링하는 통합 생성 방법을 제안하여 작업 상호 작용을 향상시키고 강력한 생성 기능을 활용하기 위해 대규모 언어 모델과의 원활한 통합을 가능하게 합니다. 공개 SLU 데이터 세트에 대한 광범위한 실험을 통해 제안된 방법의 효과를 입증하고 여러 벤치마크 방법에 비해 우수한 SLU 성능을 달성함을 보여줍니다. 모든 코드와 모델을 GitHub에 공개하여 향후 연구를 촉진할 예정입니다.

시사점, 한계점

•

시사점:

◦

단일 아키텍처에서 여러 SLU 작업을 통합적으로 모델링함으로써 시스템 복잡성을 줄이고 작업 간 상호 작용을 향상시켰습니다.

◦

이종 데이터 세트를 효율적으로 활용하여 SLU 성능을 향상시켰습니다.

◦

대규모 언어 모델과의 통합을 통해 생성 능력을 강화했습니다.

◦

우수한 SLU 성능을 달성하여 실제 음성 기반 멀티미디어 시나리오에 적합합니다.

◦

공개된 코드와 모델을 통해 향후 연구를 촉진합니다.

•

한계점:

◦

제안된 방법의 일반화 성능에 대한 추가적인 평가가 필요합니다.

◦

다양한 음성 언어 및 액센트에 대한 로버스트성을 평가할 필요가 있습니다.

◦

실제 응용 프로그램에 적용하기 위한 추가적인 연구가 필요합니다.

Made with Slashpage