Angel Sevillano

@asevillano · User

GitHub profile ↗ · Compare

2 followers27 repositories

Repositories

asevillano/genesys-voice-live-connector

This repository provides a Voice AI Agent implementation that integrates with Genesys Cloud Audio Connector using the Azure Voice Live API (Real-time Speech-to-Speech) with Azure Speech voices

★ 0TypeScriptForks 0

asevillano/stt-speech-llm

Real-time speech-to-text on Azure Speech with per-phrase intent, sentiment, and running summary via Azure OpenAI model, with optional CSV export

★ 0PythonForks 0

asevillano/Genesys_Cloud_Agent_Assist

End-to-end Agent Assist demo that mimics how the 'Active Listening' solution would integrate with Genesys Cloud, but without needing a Genesys license: a built-in web simulator plays the role of the Genesys agent desktop and the AudioHook v2 audio connector.

★ 0PythonForks 0

asevillano/foundry-vnet-deploy

This Skill deploys Azure AI Foundry with Agent Setup in a private VNet. Generates the .bicepparam file and runs the deployment with az deployment group create. Supports new or existing VNets, existing resources (CosmosDB, Storage, AI Search) and existing private DNS zones.

★ 1BicepForks 0

asevillano/aoai-sign-language-video-analysis

This project demonstrates how to use **Azure OpenAI GPT-4.1** multimodal capabilities to **recognize British Sign Language (BSL) concepts from video**, using a **few-shot prompting** approach.

★ 0Jupyter NotebookForks 0

asevillano/video-analysis-with-aoai

This repository showcases how to leverage the capabilities of LLMs to analyze and extract insights from video files or video URLs, including their audio content, offering several configurable parameters, such as the duration for splitting the video, the number of frames to extract per second, frame resizing, and prompts.

★ 0PythonForks 1

asevillano/active-listening-demo

This demo integrates the Azure Speech service (in connected or disconnected containers) to transcribe voice and synthetize agent's response, Azure AI Language (also in connected or disconnected containers) to mask PII, and an AI Foundry Agent with a knowledge base to provide answers and suggestions to the human agent during a customer conversation.

★ 0PythonForks 0

asevillano/visual-ai-search

Image search application with dual vectorization comparing Azure AI Vision multimodal embeddings vs Azure OpenAI text embeddings side-by-side. Uploaded images are analyzed by GPT-4.1 Vision with a fully customizable prompt, enabling domain-specific metadata extraction.

★ 0PythonForks 0

asevillano/voice_live

Demo of Azure Voice Live API with several options: 1) voice-live.py to speak with the model, 2) voice-live-function-calling.py adding with function calling, 3) voice-live-agents.py with a Foundry agent, and 4) voice-live-websocket-server+websocket-audio-client-mic with client/server architecture

★ 0PythonForks 0

asevillano/snippy

🧩 Demo: Build AI-powered MCP Tools for GitHub Copilot using Azure Functions, OpenAI, AI Agents & Cosmos DB. Manages code snippets w/ vector search.

★ 0Forks 0

asevillano/stt-llm-tts-demo

This repo includes a STT + AOAI + TTS using Azure OpenAI transcribe and tts models, and Azure Speech service for stt and tts

★ 0PythonForks 0

asevillano/web_crawler

This script crawls a starting URL using Selenium in headless mode, follows links up to a maximum depth (or infinitely if 0), and downloads files with the specified file extensions, and more options

★ 0PythonForks 1

asevillano/multimodal-rag-code-execution

A multimodal Retrieval Augmented Generation with code execution capabilities. Process multiple complex documents with images, table, charts to distill insights or generate new documents.

★ 0Forks 0

asevillano/dialog-tool

Tool to create, manage, and interact with dialogs for the IBM Watson Dialog Service

★ 0CSSForks 0