PyTorch implementation of "Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss" (ICASSP 2020)
-
Updated
Feb 27, 2022 - Python
PyTorch implementation of "Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss" (ICASSP 2020)
A curated list of awesome papers on contextualizing E2E ASR outputs
An efficient implementation of RNN-T Prefix Beam Search in C++/CUDA.
An implementation of RNN-Transducer loss in TF-2.0.
FunASR实时语音识别版,识别麦克风和电脑内播放的声音,电脑语音打字软件
I'm building an end-to-end Vietnamese Speech Recognition System. I'll deploy it into production with the help of Flask, Uwsgi, Nginx, and AWS ...
NVIDIA Nemotron 3.5 ASR (multilingual streaming, 600M) exported to ONNX: cache-aware encoder/decoder/joiner + a numpy+onnxruntime streaming engine. WER 0.0137 vs original, 3.8x realtime on CPU.
Pure PyTorch implementation of the loss described in "Online Segment to Segment Neural Transduction" https://arxiv.org/abs/1609.08194
🔊 Enhance speech recognition with GLM-ASR-Nano-2512, a high-performance model excelling in dialect support and low-volume audio accuracy.
Local-first Vietnamese/English streaming speech-to-text PWA with private voice personalization
Deep learning-based subtitle generation model that processes audio datasets to generate accurate text transcriptions. Includes audio feature extraction, encoder-decoder architecture, training pipelines, and evaluation metrics for subtitle alignment.
NVIDIA Conformer-CTC and Conformer-Transducer (RNN-T) running natively on Apple Silicon via MLX. Loads NeMo checkpoints directly.
🚀 Create and manage SPL tokens on the Solana blockchain with ease, using our Next.js-based launchpad for streamlined token and liquidity management.
To associate your repository with the rnnt topic, visit your repo's landing page and select "manage topics."