Ramit Pahwa

I'm a Senior Research Engineer at Rivian and Volkswagen Group Technologies in Palo Alto, working on speech, multimodal AI and post-training. I lead post-training of a 30B omni foundation model for in-vehicle spoken dialogue, and have worked on full-duplex speech models with native tool calling, on-device wake word detection, streaming multilingual ASR, and production text-to-speech.

Before that I was a Machine Learning Scientist at Secta.ai, working on diffusion models and LLM-based image quality assessment, and a Software Engineer (ML) at Adobe on Acrobat's mobile ML platform. I did my M.S. in Computer Science at UT Austin, where I worked on video quality assessment in the LIVE Lab with Prof. Alan Bovik, and my integrated B.S./M.S. in Mathematics and Computing at IIT Kharagpur.

Email  /  Resume  /  Scholar  /  Twitter  /  Github  /  LinkedIn

profile photo

Research

I work on speech and multimodal AI: post-training voice-language models for spoken dialogue, evaluating what they actually hear, and making speech models small enough to run on-device. Earlier I worked on perceptual video quality assessment, and on compressing and distilling neural networks.

VLME

VoiceLongMemEval: Do Assistants Remember How You Sounded?
Ramit Pahwa*, Parivesh Priye*, Apoorva Beedu  (* equal contribution)
arXiv, 2026  (submitted to the PALM Workshop @ NeurIPS 2026)
arXiv

A long-horizon conversational memory benchmark in which every answer depends on paralinguistic metadata (emotion, prosody, voice events) attached to the turns, filtered by a three-stage adversarial gate so that a strong language model fails given the transcript alone. Frontier models show a pervasive affect gap: adding the paralinguistic track lifts accuracy by 0.09 to 0.38, while standard ASR pipelines discard the signal entirely.

Audio2Tool

Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Ramit Pahwa*, Apoorva Beedu*, Parivesh Priye, Rutu Gandhi, Saloni Takawale, Aruna Baijal, Zengli Yang  (* equal contribution)
Interspeech, 2026  (submitted)
project page / arXiv / dataset / code

About 30,000 spoken tool-calling queries across smart car, smart home and wearables, organized in a reasoning hierarchy from single-intent calls to needle-in-a-haystack extraction, with realistic acoustic variation from zero-shot voice-cloning TTS and noise. SpeechLM and ASR-plus-LLM baselines expose the gap between text and audio tool calling.

Frames from the LIVE-ASL database LIVE-ASL: Subjective and Objective Quality Assessment of American Sign Language Videos
Sandeep Mishra, Shashank Gupta, Ramit Pahwa, Margaret H. Pinson, Alan C. Bovik
IEEE Transactions on Image Processing, 2026
project page / database

A public video quality database of 358 compression-distorted American Sign Language videos, built from 149 signed messages by 21 signers and rated for both intelligibility and quality by 42 Deaf and Hard of Hearing volunteers. Existing VQA models leave significant room for improvement on sign language video.

Video quality assessment taxonomy Perceptual Video Quality Assessment: The Journey Continues!
Avinab Saha*, Sai Karthikey Pentapati, Zaixi Shang, Ramit Pahwa*, Bowen Chen, Hakan Emre Gedik, Sandeep Mishra, Alan C. Bovik  (* equal contribution)
Frontiers in Signal Processing, 2023
paper

A survey of perceptual video quality assessment: full-, reduced- and no-reference models, the subjective databases behind them, and emerging topics such as HDR, high frame rate, cloud gaming and AR/VR.

Model Blending Model Blending for Text Classification
Ramit Pahwa, Surya Dwivedi, Jayanta Mukhopadhyay, Sunav Choudhary, Vishwa Vinay
arXiv, 2022  (Master's Thesis, IIT Kharagpur & Adobe Research)
arXiv / paper / thesis

Blending in complementary knowledge from an LSTM teacher lets us train CNN students for text classification that are up to 15x faster at inference, while being more accurate than the same CNN trained directly on the data.

Data-Driven Compression Data-Driven Compression of Convolutional Neural Networks
Ramit Pahwa, Manoj Ghuhan Arivazhagan, Ankur Garg, Siddarth Krishnamoorthy, Rohit Saxena, Sunav Choudhary
arXiv, 2019  (Adobe Research)
arXiv / code

Reinforcement learning over architecture search, combined with knowledge distillation, compresses trained CNNs such as VGG and ResNet automatically, exposing the trade-off between speed, memory and accuracy and exploiting the fact that a deployed model may only need a subset of the classes.

LSTM with attention LSTMs with Attention for Aggression Detection
Nishant Nikhil, Ramit Pahwa, Mehul Kumar Nirala, Rohan Khilnani
TRAC Workshop, COLING, 2018
paper / code

An LSTM with an attention unit for the shared task on aggression identification in Facebook posts and comments, ranking 6th and 4th on the two Hindi subtasks.

Hip ultrasound with detected femoral head Identification of Hip Dysplasia in Infants
Ramit Pahwa, Dana Cobzas
Research Internship, University of Alberta, 2017  (poster, UoA Symposium)

An end-to-end training regime for infant hip ultrasound that first localizes the femoral head with a Faster R-CNN and then segments it with a U-Net, built on a collected and cleaned set of high-dimensional MRI images.

Patents

Wake Word Detection, 2025  (filed)
Two-Stage Wake Word Detection, 2025  (filed)
Contextual Cabin Noise Suppression, 2025  (filed)
Diffusion-Based 2D Animation Creation Using Text Prompts, 2023  (approved)
Transformers for In-Application Text-Based Reading, 2021  (approved)

Design and source code from Jon Barron's website.