Just a few days after the launch of Deep Research, powered by Gemini 2.5 Pro Experimental,Google once again launches a new AI model named “DolphinGemma”This large language model is primarily used to assist scientists in researching how to communicate with dolphins, with the hope of learning what dolphins are saying.

Google launches a new AI model to help decipher dolphin language.
Google is collaborating with researchers at the Georgia Institute of Technology and the Wild Dolphin Project (WDP), whose primary mission is to observe, document, and report on the natural behaviors, social structures, communication patterns, and habitats of wild dolphins—particularly Atlantic spotted dolphins (Stenella frontalis)—through non-invasive, long-term field research. Over the years, the data collected by WDP has enabled it to correlate certain dolphin vocalizations with behaviors. According to Google, analyzing the natural and complex communication of dolphins is a formidable task, and WDP’s extensive labeled database offers a unique opportunity for AI, such as:
- The signature, one-of-a-kind whistle a mother and child use to find each other.
- The sudden croaking sound often observed during fights.
- A buzzing sound often used during courtship or when chasing sharks.

This is where DolphinGemma comes in. In short, it’s an AI model developed by Google based on the WDP database. It uses Google’s own SoundStream tokenizer to break down dolphin vocalizations into more manageable audio units. These audio units are then processed through a model architecture specifically designed to understand complex sequences. The entire setup has about 400 million parameters, yet it’s lightweight enough to run locally on a Pixel phone, making it convenient for researchers to carry around.

Unlike traditional machine learning models, DolphinGemma does not process text or images—it operates on a strict audio-in, audio-out basis. It takes in natural sequences of dolphin vocalizations, processes them using approaches inspired by how large language models understand human speech, and predicts the most likely next sound in the sequence. Dr. Denise Herzing, founder of the WDP project, has compared this to autocomplete. The model is trained to recognize patterns, structure, and progression in these sounds, just as text-based models predict the next word in a sentence based on context.
Just like other Gemma models, Google says it will bring DolphinGemma to the world this summer as an open model, hoping to provide researchers around the globe with tools to mine their own acoustic databases, accelerate the search for patterns, and collectively deepen understanding of these highly intelligent marine mammals.
Source: KOCPC Chinese