Splend1d/Zhuan
小篆學習器 learn ancient chinese characters
小篆學習器 learn ancient chinese characters
Language-Residual-Logits-Lens-Visualization-Delta
Robust Speech Recognition via Large-Scale Weak Supervision
Visualization of ASR Model logits for LLM-based an non-LLM-based approaches
Code for T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5
Repository for "Analyzing the Robustness of Unsupervised Speech Recognition", including patches to wav2vec-u and analysis code
Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.
Prompt Pool Agent
聯發創新基地(MediaTek Research) 致力於研究基礎模型。我們將研究體現在適合繁體中文使用者的模型上,並在使用權許可的情況下,提供模型給學術界研究或產業界使用。
This is a repository for Tomofun 狗音辨識 AI 百萬挑戰賽, a audio classification challenge focusing on dog sounds and noises inside the house.
Ongoing research training transformer language models at scale, including: BERT & GPT-2
Code for ACL 2022 Conference Paper "XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding"
TFDS is a collection of datasets ready to use with TensorFlow, Jax, ...
Self-Supervised Speech Pre-training and Representation Learning Toolkit.
A C++ code which uses boost library to solve the rectilinear polygon splitting problem
DUAL with run_squad
Textless (ASR-transcript free) Spoken Question Answering. The official release of NMSQA dataset and the implementation of "DUAL: Textless Spoken Question Answering with Speech Discrete Unit Adaptive Learning" paper.
Explore different way to mix speech model(wav2vec2, hubert) and nlp model(BART,T5,GPT) together
Multiple Assignments of the Course : "Theory of Computer Games"
codes and other refs for posts on medium
PyTorch code for EMNLP 2019 paper "LXMERT: Learning Cross-Modality Encoder Representations from Transformers".
Implementations of recent research prototypes/demonstrations using MONAI.