Graduate student at Carnegie Mellon University
Repositories
satvik-dixit/CPP
Python implementation of Cepstral Peak Prominence (CPP)
satvik-dixit/aura
satvik-dixit/MFCon
Code for the paper: Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
satvik-dixit/mace
Code for the paper: MACE: Leveraging Audio for Evaluating Audio Captioning Systems
satvik-dixit/EzAudio
High-quality Text-to-Audio Generation with Efficient Diffusion Transformer
satvik-dixit/versa
satvik-dixit/cmu-mlsp.github.io
[website] CMU MLSP group
satvik-dixit/explainability_SER
Code for the paper: Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
satvik-dixit/fense
Fluency ENhanced Sentence-bert Evaluation (FENSE), metric for audio caption evaluation. And Benchmark dataset AudioCaps-Eval, Clotho-Eval.
satvik-dixit/T-FOLEY
Implementation of the paper, T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis, accepted in 2024 ICASSP
satvik-dixit/AudioLDM
AudioLDM: Generate speech, sound effects, music and beyond, with text.
satvik-dixit/stable-audio-tools
Generative models for conditional audio generation
satvik-dixit/SpeechTokenizer
This is the code for the SpeechTokenizer presented in the SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models. Samples are presented on
satvik-dixit/speech_emotion_recognition
Identifying emotions from speech using self supervised learning
satvik-dixit/espnet
End-to-End Speech Processing Toolkit
satvik-dixit/ML_Forex_Forecasting
Implementation of numerous machine learning methods for the prediction of daily current exchange rates.
satvik-dixit/feature_importance
Using feature importance for speech emotion recognition
satvik-dixit/pyroomacoustics
Pyroomacoustics is a package for audio signal processing for indoor applications. It was developed as a fast prototyping platform for beamforming algorithms in indoor scenarios.