[공지사항]을 빙자한 안부와 근황

Show more

Daily Arxiv

전 세계에서 발간되는 인공지능 관련 논문을 정리하는 페이지 입니다.
본 페이지는 Google Gemini를 활용해 요약 정리하며, 비영리로 운영 됩니다.
논문에 대한 저작권은 저자 및 해당 기관에 있으며, 공유 시 출처만 명기하면 됩니다.

SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding

Romance, Relief, and Regret: Teen Narratives of Chatbot Overreliance

Supernova: Achieving More with Less in Transformer Architectures

GR-3 Technical Report

EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Contro

AI-Enhanced Precision in Sport Taekwondo: Increasing Fairness, Speed, and Trust in Competition (FST.ai)

Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models

NeuroHD-RA: Neural-distilled Hyperdimensional Model with Rhythm Alignment

IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning

Lessons from the TREC Plain Language Adaptation of Biomedical Abstracts (PLABA) track

Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters

Physical models realizing the transformer architecture of large language models

GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities

Multimodal Coordinated Online Behavior: Trade-offs and Strategies

A Survey of Deep Learning for Geometry Problem Solving

Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models

Pre-Training LLMs on a budget: A comparison of three optimizers

OPC: One-Point-Contraction Unlearning Toward Deep Feature Forgetting

Adaptive Gaussian Mixture Models-based Anomaly Detection for under-constrained Cable-Driven Parallel Robots

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving

Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling

The Joys of Categorical Conformal Prediction

Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on Edge

Capacity Planning and Scheduling for Jobs with Uncertainty in Resource Usage and Duration

FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization

Neural Approaches for Multi-Objective Routing on Multigraphs

Diffusion-Based Electrocardiography Noise Quantification via Anomaly Detection

CogStream: Context-guided Streaming Video Question Answering

Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems

Towards provable probabilistic safety for scalable embodied AI systems

SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods

Human Empathy as Encoder: AI-Assisted Depression Assessment in Special Education

Multimodal Forecasting of Sparse Intraoperative Hypotension Events Powered by Language Model

Autocomp: LLM-Driven Code Optimization for Tensor Accelerators

ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection

ReMi: A Random Recurrent Neural Network Approach to Music Production

Aitomia: Your Intelligent Assistant for AI-Driven Atomistic and Quantum Chemical Simulations

A Goal-Oriented Reinforcement Learning-Based Path Planning Algorithm for Modular Self-Reconfigurable Satellites

A Method for the Architecture of a Medical Vertical Large Language Model Based on Deepseek R1

Balancing Robustness and Efficiency in Embedded DNNs Through Activation Function Selection

Antithetic Sampling for Top-k Shapley Identification

GeoFlow-SLAM: A Robust Tightly-Coupled RGBD-Inertial and Legged Odometry Fusion SLAM for Dynamic Legged Robotics

SciFi-Benchmark: Leveraging Science Fiction To Improve Robot Behavior

ETCH: Generalizing Body Fitting to Clothed Humans via Equivariant Tightness

Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $\mu$P Parametrization

Curating Demonstrations using Online Experience

OMNISEC: LLM-Driven Provenance-based Intrusion Detection via Retrieval-Augmented Behavior Prompting

PRISM: High-Resolution & Precise Counterfactual Medical Image Generation using Language-guided Stable Diffusion

Reasoning Does Not Necessarily Improve Role-Playing Ability

BioMaze: Benchmarking and Enhancing Large Language Models for Biological Pathway Reasoning

Revealing Bias Formation in Deep Neural Networks Through the Geometric Mechanisms of Human Visual Decoupling

Conformal Predictions for Human Action Recognition with Vision-Language Models

Cross-Encoder Rediscovers a Semantic Variant of BM25

DisCoPatch: Taming Adversarially-driven Batch Statistics for Improved Out-of-Distribution Detection

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

R-Bot: An LLM-based Query Rewrite System

Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation

Atomic Calibration of LLMs in Long-Form Generations

LibEER: A Comprehensive Benchmark and Algorithm Library for EEG-based Emotion Recognition

Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis

VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models

From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent AI safety benchmarks

V-RoAst: Visual Road Assessment. Can VLM be a Road Safety Assessor Using the iRAP Standard?

FLAIN: Mitigating Backdoor Attacks in Federated Learning via Flipping Weight Updates of Low-Activation Input Neurons

FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image Translation

Analysis of the 2024 BraTS Meningioma Radiotherapy Planning Automated Segmentation Challenge

Unisolver: PDE-Conditional Transformers Are Universal PDE Solvers

LangBiTe: A Platform for Testing Bias in Large Language Models

Practical Insights into Knowledge Distillation for Pre-Trained Models

Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy

Energy-Efficient and Real-Time Sensing for Federated Continual Learning via Sample-Driven Control

Gemini 2.5 Pro Capable of Winning Gold at IMO 2025

Hierarchical Budget Policy Optimization for Adaptive Reasoning

BioGraphFusion: Graph Knowledge Embedding for Biological Completion and Reasoning

Routine: A Structural Planning Framework for LLM Agent System in Enterprise

CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning

Assessing Adaptive World Models in Machines with Novel Games

A Multi-granularity Concept Sparse Activation and Hierarchical Knowledge Graph Fusion Framework for Rare Disease Diagnosis

An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis

Hierarchical Reasoning Model

The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning

DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph

InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification

R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory

Efficient Strategy Learning by Decoupling Searching and Pathfinding for Object Navigation

Alto: Orchestrating Distributed Compound AI Systems with Nested Ancestry

Toward A Causal Framework for Modeling Perception

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Decoding Translation-Related Functional Sequences in 5'UTRs Using Interpretable Deep Learning Models

Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning

Never Come Up Empty: Adaptive HyDE Retrieval for Improving LLM Developer Support

AI-enhanced conversational agents for personalized asthma support Factors for engagement, value and efficacy

RAVine: Reality-Aligned Evaluation for Agentic Search

Experience is the Best Teacher: Grounding VLMs for Robotics through Self-Generated Memory

Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy

Created by

Haebom

저자

Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, Arman Cohan, Zhiyong Lu, Mark Gerstein

개요

본 논문은 대규모 언어 모델 기반의 AI 과학자의 잠재적 위험성을 다룬다. AI 과학자는 다양한 분야에서 실험을 자율적으로 수행하고 과학적 발견을 촉진하는 데 상당한 가능성을 보여주지만, 동시에 새로운 취약성을 야기한다. 본 논문은 사용자 의도, 특정 과학 분야, 외부 환경에 미치는 잠재적 영향 등을 고려하여 AI 과학자의 잠재적 위험을 개괄하고, 이러한 취약성의 근본 원인을 탐구하며 기존 연구들을 검토한다. 나아가, 인간 규제, 에이전트 정렬, 환경 피드백 이해(에이전트 규제)를 포함하는 3중 구조 프레임워크를 제안하여 위험을 완화하고, 향상된 모델, 강력한 벤치마크, 포괄적인 규정 개발의 필요성을 강조한다.

시사점, 한계점

•

시사점: AI 과학자의 위험성에 대한 포괄적인 분석을 제공하고, 인간 규제, 에이전트 정렬, 환경 피드백 이해를 통합한 위험 완화 프레임워크를 제시한다. 향상된 모델, 강력한 벤치마크, 포괄적인 규정의 중요성을 강조한다.

•

한계점: AI 과학자의 위험성에 대한 연구가 아직 초기 단계이며, 제시된 프레임워크의 실효성 검증이 필요하다. AI 과학자의 발전 속도를 고려할 때, 규제 및 안전 조치의 개발과 적용이 시급한 과제로 남는다.

Made with Slashpage