배치된 시스템을 위한 에이전트 수명 엔지니어링
인공지능(AI) agent가 장기적으로 운영되면서 운영 시스템으로 사용되면서, 여전히 초기화된 모델처럼 평가되고 있다. 첫 번째 운영일(day-one) 기준점은 시스템의 기본적인 질문을 놓치고 있다. 그 질문은 agent가 배포된 후 얼마 동안 신뢰할 수 있는 상태를 유지하는지다. 모델의 가중치가 고정되어 있어도, agent의 실제 상태는 계속해서 변하고 있다. 이는 agent가 상호 작용
1186개의 기사
인공지능(AI) agent가 장기적으로 운영되면서 운영 시스템으로 사용되면서, 여전히 초기화된 모델처럼 평가되고 있다. 첫 번째 운영일(day-one) 기준점은 시스템의 기본적인 질문을 놓치고 있다. 그 질문은 agent가 배포된 후 얼마 동안 신뢰할 수 있는 상태를 유지하는지다. 모델의 가중치가 고정되어 있어도, agent의 실제 상태는 계속해서 변하고 있다. 이는 agent가 상호 작용
현재 도메인 지식 artifact에서 수학적 프로그래밍 모델(model)로의 제약 학습(Constraint Acquisition, CA) 및 관련 연구는 충분한 기준점의 부족으로 제한되고 있습니다. 이러한 결함은 재현성과 연구간 비교의 어려움으로 CA 방법의 성숙을 늦추고 있습니다. 기존의 기준점은 솔버 평가를 위한 것이었지만 CA 알고리즘을 평가하기 위한 것이었습니다. 또한 기준점은 구체적
대규모 멀티모달 언어 모델(LLLM)- 기반 몸체(agent)가 물리 환경에서 복잡한 작업을 해결하는 데 강력한 잠재력을 보인다. 하지만 개인화된 지원은 일반적인 지시를 따르거나 물체 카테고리만 식별하는 것보다 더 필요하다. 실제 세계 시나리오에서 목표는 종종 이전 상호 작용을 통해 암시적으로만 지정되며, 시간이 지남에 따라 축적된 개인화된 맥락을 활용하여 agent가 개인화된 맥락과 시각적
장기 AI agent는 지속적인 메모리가 필요합니다. 메모리는 학습을 위해 세션을跨越하고 반복적인 컨텍스트 주입을 줄이고 과거의 결정의 감사성을 지원합니다. 그러나 현재의 agent 메모리 시스템과 데이터베이스 패러다임은 메모리를 저장소로 취급하고, 기록, 임베딩, 또는 에지의 정확성을 localize합니다. 그러나 장기 메모리의 요구 사항을 충족하는 것은 불가능합니다. 따라서 장기 AI agent 메모리의 데이터 기초를 다시 생각할 필요성이 있습니다.
대규모 언어 모델(LLM)이 내부 상태를 감지하고 보고할 수 있는지 여부에 대한 연구가 여러 차례 진행되었지만, 이러한 결론이 미리 맏기기에 지나치다고 생각합니다. 인간의 자기 인식 연구에서 배운 교훈을 바탕으로, 내부 상태를 진정하게 인식하는 것과 표면적인 자극에 기반한 패턴 매칭을 구별하는 것은 필수적이라고 생각합니다. 또한 단지 행동적 증거만으로는 강력한 자기 인식 주장을establis
3D 형태를 기반으로 물리적으로 구축할 수 있는 벽돌 구조를 생성하려면 단순한 기하학적 복원만 아니라, 분할된 부품 제약 조건과 구조적 안정성도 충족해야 한다. 현재의 벽돌 생성 방법은 대개 힌트 기반 최적화에 의존하는데, 목표 3D 형태가 정의된 제약 조건하에서 가능성이 있는 구조를 허용하지 않을 경우 최적화가 붕괴할 수 있다. 또는, 벽돌 시퀀스를 생성하는 데 3D 기하학 및 조립 관계를
arXiv:2605.23940v1 Announce Type: new Abstract: How do multi-turn reasoning systems fail? The expected answer is logical contradiction, in which the system's maintained state becomes unsatisfiable. We
arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different t
arXiv:2605.23938v1 Announce Type: new Abstract: Large language models (LLMs) increasingly fuse heterogeneous inputs in ubiquitous systems. Yet, how LLMs implicitly allocate authority when sensor measu
arXiv:2605.23937v1 Announce Type: new Abstract: Knowledge base (KB) embeddings aim at combining the capability of classical knowledge graph embeddings to generalize the information present in facts, t
arXiv:2605.23936v1 Announce Type: new Abstract: This book presents a comprehensive and systematic survey of graph theory under uncertainty, with particular emphasis on the unifying role of the uncerta
arXiv:2605.23935v1 Announce Type: new Abstract: Autonomous agent systems fail not only due to incorrect decisions, but due to executing decisions whose authority no longer holds at runtime. Prior work
arXiv:2605.23934v1 Announce Type: new Abstract: Quantum computing devices are recognized as powerful tools for solving NP-complete problems. However, the intricacy of their modeling presents notable b
arXiv:2605.23932v1 Announce Type: new Abstract: Despite strong medical benchmark accuracy, LLMs can exhibit severe multi-turn sycophancy in clinical dialogue, abandoning initial correct diagnosis unde
arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp
arXiv:2605.23930v1 Announce Type: new Abstract: We introduce \emph{Quantum Frog}, a two-player cooperative game built on a novel \emph{quantized-time} mechanic in which the environment advances only w
arXiv:2605.23929v1 Announce Type: new Abstract: Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by con
arXiv:2605.23928v1 Announce Type: new Abstract: We present Context, the intelligence layer of the Magarshak Architecture, which replaces reactive query-response chatbots with proactive goal-directed a
arXiv:2605.23926v1 Announce Type: new Abstract: Reasoning-capable large language models solve hard problems by emitting long chains of thought, paying heavily in latency, GPU time, and energy. Casual
arXiv:2605.23909v1 Announce Type: new Abstract: We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the c
arXiv:2605.23908v1 Announce Type: new Abstract: We are in the midst of large-scale industrial and academic efforts to automate the processes of scientific, technological and creative production throug
arXiv:2605.23238v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior i
arXiv:2605.23218v1 Announce Type: new Abstract: Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly
arXiv:2605.23204v1 Announce Type: new Abstract: Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding,