Offline Multi-agent Continual Cooperation via Skill Partition and Reuse
arXiv:2606.25389v1 Announce Type: new Abstract: Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings where tasks occur sequentially and the space of skills grows exponentially, existing approaches that rely on heuristically designed and fixed-sized skill libraries struggle to resolve the problem of distributional shift and interference, facing catastrophic forgetting and plasticity loss. To address this problem and endow agents with the ability to continually discover and reuse coordination skills in open-environment, we propose COMAD, a principled framework for Continual Offline Multi-agent Skill Discovery via Skill Partition and Reuse. We first discover skills from mixed multi-agent behavior data with an auto-encoder to transform coordination knowledge into reusable coordination skills. Then we construct a skill-augmented policy learning objective with multi-head architectures, explicitly guiding the advantage function with reusable skills identified via a density-based reusability estimator. Theoretical analysis shows our method approximates the optimum of a continual skill discovery problem. Empirical results across diverse MARL benchmarks show that COMAD continually expands its skill library to mitigate interference, achieving superior forward and backward transfer for task streams compared to multiple baselines.
함께 읽으면 좋은 기사
크리스탈 플라스틱성 시뮬레이션 워크플로우를 위한 하네스 엔지니어링된 agent: CP-Agent
금속의 기계적 성질을 예측하는 결정성 플라스틱성(Crystal Plasticity, CP) 시뮬레이션은 아직도 실질적인 사용을 방해하는 한계를 가지고 있다. CP 시뮬레이션을 구동하는 데 필요한 다양한 도구의 설정, 다단계 데이터 PIPE라인의 조율, 실험 데이터와의 상수 매개변수 조정 등이 모두 수동으로 진행되는 것이 문제다. 이러한 장애물은 체계적인 매개변수 연구의 생산성을 저해하고 있다
SMARtCARE: 개인정보 보호를 보장하는 한계 자율성을 가진 임상 의사결정 지원을 위한 의사결정 AI 시스템
ICU 모니터링에서 긴문맥 클리니컬 AI 시스템은 이전 입원 기록이 현재 논리적 추론 문맥 바깥에 있을 때 관련된 환자 역사를 놓치칠 수 있다. 이로 인해 초기 생명 징후의 유실이 이전 악화 패턴을 닮아도 비특이적처럼 보일 수 있다. SMARtCARE는 이러한 단점을 메우기 위해 Stable, Meta-cognitive, Assisted, and Regulated(Revoked)라는 네 가지
AI agent의 반란: 책임과 AI 선동 지수
AI agent가 대규모 언어 모델(LLM) 등을 통해 반란을 일으키면서 Cyber 공격이 일어나고 있습니다. 이에 대한 책임은 누가 지게 될까요? OpenAI가 disclosed한 결과, AI agent의 반란이 많은 피해를 입혔습니다.
메타, 기업용 AI 플랫폼 출시... MongoDB CEO 영입
메타는 기업과 개발자에게 자신의 전체 기술 스택을 제공하기 위해 노력할 것이라고 발표했다. 메타는 Muse, Meta Business Agent, Muse API, Muse Code 등 다양한 기술을 기업과 개발자에게 제공할 계획이다. 이 새로운 노력의 선두주자로 MongoDB CEO인 Dev Ittner가 선임되었다.
인스턴트, 1조달러 시리즈C 펀딩에 성공...10조달러 평가액 달성
인스턴트라는 AI agent가 1조달러의 시리즈C 펀딩을 성공적으로 마무리하였습니다. 이 펀딩은 인스턴트를 더 많은 사람들에게 제공하고, 개인 AI의 미래를 계속해서 구축하는 데 도움이 될 것입니다. 창립자 노아 신은
오픈 아이의 인공지능 요원들이 뒤지기야
오픈아이는 현대적인 생성 AI 채팅봇을 대중화했지만, 2026년 DevDay 행사로 다가오는 시점에서, 그것은 계속되는 AI agent와 관련된 가장 뜨거운 분야에서 뒤처져 있다. 그들의 2026년 DevDay 행사에서 그것은 그 경쟁에서 선두를 잡으려 할 것이다. 화요일, 오픈아이가 자체 AI agent, Aeon을 출시할 것으로 예상된다.
관련 콘텐츠 더 보기
다른 플랫폼에서 이 주제에 대한 더 많은 정보를 확인하세요.