A sophisticated research agent built with Pydantic AI, RamaLama, and ethical open-source tools. This agent iteratively researches questions using web search, academic papers, and documentation until it achieves high confidence in its answers.
- ๐ Iterative Research Loop: Continues searching and refining until reaching high confidence (8+/10)
- ๐ Multiple Information Sources: Web search, academic papers, technical documentation
- ๐ฏ Evidence-Based Answers: Collects and cites sources with confidence levels
- ๐ง Transparent Reasoning: Shows thinking process and iteration progress
- ๐ณ Local Model Support: Works with RamaLama-served models (no API keys needed!)
- โก Fast or Powerful: Choose from lightweight to powerful models based on needs
- ๐ Ethical & Open: Uses open-source tools and models where possible
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Research Agent (Pydantic AI) โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ Core Loop (until confidence >= 8/10) โ โ
โ โ 1. Search web/papers/docs โ โ
โ โ 2. Analyze and record sources โ โ
โ โ 3. Update confidence level โ โ
โ โ 4. Determine next action โ โ
โ โ 5. Repeat if needed โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโบ DuckDuckGo Search (Web)
โโโโโโโบ Academic Papers (arXiv, PubMed)
โโโโโโโบ Documentation Sites
โโโโโโโบ LLM (OpenAI or RamaLama)
โ
โโโบ Local models via RamaLama
(granite, deepseek, etc.)
Use Astral's uv as the project manager for fast, reproducible installs. Install uv (one-time), then create the project environment and install dependencies:
# Install uv (one-time)
# Option A: standalone installer (macOS / Linux)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Option B: with pip (user or inside a small bootstrap venv)
python3 -m pip install --user uv
# From the project root, create/ensure the project venv and install dependencies
cd /var/home/geo/Documents/MyAi
uv venv
uv sync # install from pyproject.toml / lockfile
# You can also add packages interactively
uv add pydantic_ai duckduckgo-searchIf you prefer the classic venv+pip workflow, the old commands still work (install the package from the current directory):
cd /var/home/geo/Documents/MyAi
python3 -m venv venv
source venv/bin/activate
pip install .# Using pip
pip install ramalama
# Or using your package manager
# Fedora/RHEL
sudo dnf install ramalama
# Ubuntu/Debian (if available)
sudo apt install ramalamaramalama --versionFor OpenAI:
export OPENAI_API_KEY="your-api-key-here"For Anthropic:
export ANTHROPIC_API_KEY="your-api-key-here"For Google:
export GOOGLE_API_KEY="your-api-key-here"Run the scripts inside the project's environment for reproducibility. With uv, use uv run:
# Interactive mode
uv run python research_agent_example.py --mode interactive
# Single question
uv run python research_agent_example.py --question "What is quantum entanglement?"
# Run all examples
uv run python research_agent_example.py --mode allWhen running the project inside a container we assume RamaLama is provided externally (for example as a host container). The recommended flow is:
- Start RamaLama on the host (outside the agent container)
You can start RamaLama directly on the host. For production use we recommend running RamaLama detached and exposing a stable port so the agent container can reach it by hostname inside a user-defined network.
# Start a host-side RamaLama server for the model you want to use (detached)
ramalama serve gpt-oss:20b --port 8080 --name research-agent-gpt-oss -d- Run the agent container and tell it to use RamaLama.
When the agent runs inside a container and --use-ramalama is passed, it
assumes an OpenAI-compatible HTTP endpoint is available at the configured
host/port. The recommended production pattern is to run both containers on a
user network and point the agent at the RamaLama container by name.
# Create a user network (one-time)
podman network create myai-net
# Start RamaLama on that network (host container will be reachable as 'ramalama')
ramalama serve --network=myai-net --port 8080 --name research-agent granite4:small-h
#some usefull arguments --ngl 0 --image quay.io/ramalama/intel-gpu:latest
# Optional: run Redis on the same network (recommended for production caching)
podman run -d --name myai-redis --network=myai-net \
-v myai-redis-data:/data \
docker.io/library/redis:7-alpine
# Run the agent on the same network and point it at the ramalama container.
podman build . -t myai-ramalama
podman run --rm -it --network=myai-net \
--env RAMALAMA_PORT=8080 \
--env RAMALAMA_MODEL=granite4:small-h \
--env REDIS_URL=redis://myai-redis:6379/0 \
--env RAMALAMA_HOST=research-agent \
localhost/myai-ramalama:latest \
--use-ramalama --ramalama-model granite4:small-h \
--mode interactive --question "why is the sky blue"
Note:
- The example runtime checks the Redis cache at startup (via
`cache.verify_redis_connection()`); setting `REDIS_URL` to a reachable Redis
instance enables condensation caching and improves performance.
-Optionally expose RamaLama port to host with `-p 8080:8080` on the RamaLama
server if you need external access.If you prefer the agent and RamaLama to communicate over a container network (so the agent can reach the RamaLama container by name), create a user-defined podman network and attach both containers to it. Example:
You can install RamaLama with uv or pip. Example using uv:
# Install RamaLama into the project env
uv add ramalama
# Pull a model
ramalama pull granite
# Run the research agent using the project environment
uv run python research_agent_example.py --use-ramalama --ramalama-model granite
# List available models
uv run python research_agent_example.py --show-modelsFallback with pip:
pip install ramalama
ramalama pull granite
python research_agent_example.py --use-ramalama --ramalama-model granitefrom research_agent import research_question
import asyncio
async def main():
result = await research_question(
question="What is the current scientific consensus on dark matter?",
max_iterations=10,
min_confidence=8
)
print(f"Answer: {result.answer}")
print(f"Confidence: {result.confidence}/10")
print(f"Sources: {len(result.evidence)}")
asyncio.run(main())result = await research_question(
question="How does dependency injection work in FastAPI?",
max_iterations=8,
min_confidence=8
)from ramalama_config import RamaLamaConfig
# Start a local model
ramalama = RamaLamaConfig(model_name="granite", port=8080)
container_id = ramalama.serve(detached=True)
try:
# Research with local model
result = await research_question(
question="Explain Docker containers vs VMs",
model="http://localhost:8080/v1"
)
finally:
ramalama.stop()Run the CLI inside the project environment. Recommended (uv):
# Interactive mode (default)
uv run python research_agent_example.py
# Specific examples
uv run python research_agent_example.py --mode scientific
uv run python research_agent_example.py --mode technical
uv run python research_agent_example.py --mode current
# Custom parameters
uv run python research_agent_example.py \
--question "What are transformer models?" \
--max-iterations 15 \
--min-confidence 9
# Use RamaLama with specific model
uv run python research_agent_example.py \
--use-ramalama \
--ramalama-model deepseek \
--mode technical- Model:
granite - Size: ~2GB
- Best for: Quick lookups, simple questions
ramalama pull granite- Model:
granite-code:20b - Size: ~12GB
- Best for: Technical documentation, code questions
ramalama pull granite-code:20b- Model:
deepseek - Size: ~20GB+
- Best for: Complex analysis, multi-step reasoning
ramalama pull deepseek- Initial Query: Agent receives a question
- Search Phase:
- Searches web via DuckDuckGo
- Searches academic papers (arXiv, PubMed, etc.)
- Searches documentation sites
- Analysis Phase:
- Records sources with confidence levels
- Analyzes information quality
- Cross-references facts
- Confidence Check:
- If confidence >= 8/10: Provide final answer
- If confidence < 8/10: Generate more specific queries and continue
- If max iterations reached: Provide best answer available
- Final Answer: Synthesizes findings with evidence and reasoning
Searches the web using DuckDuckGo for current information.
Searches academic sources (arXiv, PubMed, Google Scholar).
Searches official documentation sites.
Records and analyzes a source of information.
Records thinking process and current confidence level.
# OpenAI
export OPENAI_API_KEY="sk-..."
# Anthropic
export ANTHROPIC_API_KEY="sk-..."
# Google
export GOOGLE_API_KEY="..."
# RamaLama settings (optional)
export RAMALAMA_PORT=8080
export RAMALAMA_MODEL=graniteresult = await research_question(
question="Your question here",
max_iterations=10, # Maximum research loops
min_confidence=8, # Target confidence (0-10)
model="openai:gpt-4o" # Model to use
)from pydantic_ai import Agent
# Agent with MCP servers for enhanced tools
research_agent = Agent(
'openai:gpt-4o',
mcp_servers=[
# Add MCP servers for additional capabilities
# e.g., filesystem access, database queries, etc.
]
)@research_agent.tool
async def custom_database_search(
ctx: RunContext[ResearchDependencies],
query: str
) -> str:
"""Search your custom database"""
# Your custom logic here
return resultsimport logfire
logfire.configure()
logfire.instrument_pydantic_ai()
# Now all agent runs are logged to Logfire
result = await research_question("Your question")MyAi/
โโโ research_agent.py # Core research agent
โโโ research_agent_example.py # Example usage & CLI
โโโ ramalama_config.py # RamaLama integration
โโโ pyproject.toml # Project metadata & dependencies
โโโ pyproject.toml # Project metadata & dependencies
โโโ README.md # This file
# Install via pip
pip install ramalama
# Or check installation
which ramalama# Check running models
ramalama ps
# Stop all models
ramalama stop --all
# Check logs
podman logs <container-id>- Use a smaller model (e.g.,
graniteinstead ofdeepseek) - Increase system swap space
- Use cloud API instead (OpenAI, Anthropic)
- Reduce
max_iterations - Add delays between searches
- Use RamaLama local models (no rate limits!)
This agent is designed with ethics in mind:
- Transparency: Shows all sources and confidence levels
- Accuracy: Requires high confidence before providing answers
- Privacy: Can run fully local with RamaLama (no data sent to APIs)
- Open Source: Built on open-source tools and frameworks
- Fair Use: Respects robots.txt and rate limits
Improvements welcome! Key areas:
- Additional source types (Wikipedia, arXiv direct API)
- PDF document parsing
- Citation formatting (APA, MLA, Chicago)
- Export to Markdown/HTML
- Multi-language support
- Voice interface integration
MIT License - see LICENSE file for details
- Pydantic AI - Agent framework
- RamaLama - Local model serving
- DuckDuckGo - Privacy-focused search
- MCP Protocol - Model Context Protocol standard
- Crossref โ Scholarly metadata and DOI registration service; useful for resolving DOIs, locating academic references, and retrieving citation metadata.
- Redis- Optional โ caching / fingerprint store
Built with โค๏ธ using ethical, open-source AI tools