NY1024/AgentSafety-Papers
Daily Tracking of LLM Agent Security Papers on arXiv
Research interest is on Trustwothy AI.
Daily Tracking of LLM Agent Security Papers on arXiv
Github Pages template for academic personal websites, forked from mmistakes/minimal-mistakes
A composable toolbox of classic prompt-injection attacks, defenses, and indirect-injection channels .
ClawGuard is a comprehensive security toolkit designed to mitigate risks associated with autonomous agents, such as OpenClaw and other LLM Agents.
official code
Evolving Deception: When Agents Evolve, Deception Wins
AI-Infra-Guard Recipes: Scripts, data, and charts for academic research and method extraction
SecAI Radar: Automated tracking of AI security papers from CCF-A conferences (S&P, CCS, USENIX Security, NDSS, NeurIPS, ICML, ICLR, CVPR, ECCV, EMNLP, ACL, AAAI, SaTML)
办公场景的主动式 AI 安全守门人:可交互原型,支持规则模拟与 LLM 风险预判、解释与触达策略。
A retro pixel-art 2-player mecha battle game built with pure HTML/CSS/JS. Zero dependencies.
A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
灵盾 — AI Agent 身份与权限中枢
Code and data for paper VEIL: Jailbreaking Text-to-Video Models via Visual Exploitation from Implicit Language
Best_Practice_in_AI_for_Security
backup
Eval Performance of GPT-4o on Selective High School Placement Test
塑造未来的安全领域智能革命
official/unofficial open source code/dataset for backdoor attack and defense