August-murr/Interregnum
A public notebook. Working drafts, half-formed ideas, and thoughts I'm too much of a perfectionist to just post.
A public notebook. Working drafts, half-formed ideas, and thoughts I'm too much of a perfectionist to just post.
Ouroboros is a self-referential AI experimentation framework where an agent evolves the very codebase that defines its own behavior.
The repo for my submission for APARTs Secret Loyalties Hackathon
The repo for all files and reports for the Bluedot Impact AI safety project
Code and data for "Sandboxed AI Agents with GPU Access" blog post. An OpenHands agent with GPU access that completes ML tasks. Part of the Ouroboros research project.
verl: Volcano Engine Reinforcement Learning for LLMs
Lightweight OpenHands CLI in a binary executable
train agents to maximize productivity per dollar, learning to optimize tool-use, exploration,and verification under strict budgets.
Train transformer language models with reinforcement learning.
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
🤗 smolagents: a barebones library for agents. Agents write python code to call tools and orchestrate other agents.
A high-throughput and memory-efficient inference and serving engine for LLMs
Democratizing Reinforcement Learning for LLMs
LeetGPU Challenges
My AI Lab
Fully open reproduction of DeepSeek-R1
scale inference-time compute of open models