A codec is a small, strange machine: it takes sound apart, throws away (or very cleverly repackages) what you will not miss, and builds it back on the other side. Every format you have ever double-clicked is somebody's answer to the question “how little can we store and still get away with it?” — and the answers are wildly different depending on who was asking: a record label, a phone company, a game studio, a broadcaster, a hobbyist with a hard drive full of FLAC.

We have spent the last stretch teaching chainDRiVE to speak a lot of those answers, including a few we had to implement ourselves because nobody had shipped them for a browser. This is a tour of the zoo: the famous animals, the weird ones, and the ones still on our wish list.

Four ways to shrink sound

Almost every codec is a variation on one of four ideas:

  • Perceptual (transform) coding. Chop the signal into frequency bands and spend bits only where the ear is paying attention. Loud sounds mask quieter neighbours, so those can be coarsely quantized. This is MP3, AAC, Vorbis, Opus. It is lossy by design, and it is brilliant.
  • Lossless coding. Predict the next sample from the previous ones, then store only the (small) prediction error with an efficient entropy code. Decode and you get every bit back. FLAC and its relatives live here.
  • ADPCM. The humble approach: predict, store a tiny 2- to 5-bit difference, adapt the step size as you go. No psychoacoustics, almost no CPU. It is cheap, fast, and was the sound of an entire era of hardware.
  • Speech vocoders. Model the voice — a buzzing source pushed through a vocal-tract filter — and transmit the parameters instead of the waveform. Fantastic for speech at very low bitrates, hopeless for your favourite album.

The mainstream

MP3 (MPEG-1 Layer III) is the one that changed the world: a hybrid filterbank, a psychoacoustic model, and Huffman coding, tuned until “near-CD” fit in a fraction of the space. Thirty-odd years on, a good encoder like LAME is still remarkable, which is why chainDRiVE keeps all of LAME's knobs within reach. AAC is its successor, usually better at the same bitrate, and it is a whole family: plain AAC-LC, HE-AAC v1, which adds Spectral Band Replication (SBR) so the encoder transmits only the lower half of the spectrum and the decoder reconstructs the highs from a few hints, and HE-AAC v2, which adds Parametric Stereo (PS) on top, sending a mono signal plus a description of the stereo image. At very low bitrates it is dark magic.

Opus (an IETF standard) merges a speech-oriented layer with a music-oriented one and switches between them on the fly; it is what modern voice chat and a great deal of streaming run on. Ogg Vorbis was the open, patent-free answer to MP3 years earlier and is still a fine codec. FLAC is the free lossless standard; ALAC is Apple's lossless format, in the same spirit. Those six cover most of what people listen to.

Telephony, the lean years

Phone audio is a story of ruthless economy. G.711, with its two flavours A-law and μ-law, dates to the early 1970s: it squeezes 13- or 14-bit-ish dynamic range into 8 bits with a logarithmic curve, because quiet sounds need fine steps and loud ones do not. GSM 06.10 (“full rate”) went further, modelling speech with long- and short-term prediction down to about 13 kbps. G.726 is ADPCM for 8 kHz voice at 16, 24, 32 or 40 kbps — 32 being the classic “same quality as G.711 at half the bits” — and G.722 splits 16 kHz audio into two sub-bands and ADPCM-codes each, which is why “HD voice” on office phones sounds so much clearer than a regular call.

Then there is OKI/Dialogic ADPCM (the .vox files): 4-bit ADPCM at 6 or 8 kHz, a staple of voicemail and IVR boards in the computer-telephony boom. If you have ever pressed “1 for sales” on a system from the 90s, you may well have heard one. IMA ADPCM and Microsoft ADPCM are the same family wearing WAV headers, the default “small” audio on 90s PCs.

MPEG Layers I and II: the broadcaster's friend

Before Layer III there were Layers I and II. Layer II (MP2) is simpler than MP3, with no MDCT and a plainer filterbank, and that is a feature: it is robust, low-latency-friendly, and it sounds good at generous bitrates. It was chosen for DAB digital radio, widely used in digital TV broadcasting, and is the audio of Video CD; DVD-Video allows it too. It still matters in broadcasting today: plenty of radio contribution links, playout systems, and archives speak MP2, and the broadcast WAV family can carry MPEG audio inside it. Layer I (MP1) is the even simpler ancestor, used in Philips' DCC cassette, and almost nobody makes one anymore — which is exactly why it is fun to support.

The game-console crowd

Cartridge and disc space was always the enemy, and the people who wrote console audio were quietly brilliant.

  • Amiga 8SVX is an IFF file of 8-bit samples, with an optional Fibonacci-delta mode: each sample is stored as a 4-bit index into a small table of deltas shaped like the Fibonacci sequence (…, 3, 5, 8, 13, 21…). Fine steps when the signal moves slowly, big steps when it moves fast — and half the size.
  • Creative VOC is the Sound Blaster-era container: a list of typed blocks holding sample data, silences and repeats.
  • Sony PlayStation SPU-ADPCM (VAG). Audio is stored in 16-byte blocks: one byte for a predictor filter and a range shift, one for flags like loop start and end, and fourteen bytes of 4-bit codes, 28 samples per block. The filter chooses how to predict from the previous two samples (there are five of them in the console's SPU); the shift scales the residual. Encoders try each combination per block and keep the best one. That is about 3.5 times smaller than raw 16-bit PCM, with hardware doing the decoding for free.
  • PlayStation CD-ROM XA used a related ADPCM, with four predictor filters and range shifts again, to put audio on a standard CD in a form that could be interleaved with video or data. XA's 4-bit levels run at 37.8 kHz (stereo-friendly Level B) or 18.9 kHz (Level C), with an 8-bit Level A at 37.8 kHz too. Because the drive could play such sectors while it kept reading others, a 1x CD could carry streaming audio and full-motion video without a stutter.
  • Nintendo DSP-ADPCM (GameCube and Wii) packs 14 samples into every 8-byte frame, using a per-channel table of predictor coefficients stored in the file header, so the encoder can tune the predictors to the actual sound.
  • CRI ADX is the middleware codec that showed up in many Japanese and Sega titles; it is ADPCM with a fixed predictor and a configurable high-pass cutoff in the header.
  • Yamaha ADPCM was baked into sound chips such as the YMZ280B and the AICA used in the Dreamcast, another very compact 4-bit format.

Containers: broadcast, Apple and big files

A container is the envelope, and the envelope has politics. The classic WAV tops out at 4 GB and says nothing about who recorded it. The European Broadcasting Union's Broadcast Wave Format (BWF, EBU Tech 3285) adds a bext chunk — description, originator, date and time, a time reference, a coding history — so a file carries its own provenance through a broadcast chain. RF64 (EBU Tech 3306), and its ITU successor BW64, break the 4 GB barrier for long multichannel recordings. Sony's Wave64 is another 64-bit take on the same idea, and Apple's CAF (Core Audio Format) was built to hold practically anything without a size limit. Different envelopes, same job: stop the file format from being the thing that runs out first.

The lossless and audiophile cabinet

Past FLAC, there is a whole cabinet of lossless codecs. Monkey's Audio chases the smallest files with heavy, slow-to-encode adaptive filtering. WavPack famously offers a hybrid mode: a small lossy file plus a correction file that restores the original exactly. TTA (True Audio) is a fast, simple adaptive coder. TAK and OptimFROG sit at opposite ends of a speed-versus-size trade-off. Shorten, from the 1990s, was the ancestor that showed the world that linear prediction made lossless audio practical. DSD (the 1-bit format behind SACD) does not fit the pattern at all: very high-rate single-bit audio, with DST as its lossless compression.

On the lossy side of the audiophile world sits Musepack (MPC). Descended from MPEG-1 Layer II technology and developed by Frank Klemm and others, it was tuned by listening at high bitrates rather than for the lowest possible size, and it earned a loyal following for transparency. Elsewhere: ATRAC (MiniDisc, and later in some Sony devices), WMA (Microsoft's answer to MP3 and AAC), AC-3 (Dolby Digital) and DTS for surround in cinemas and on discs, Speex (a free speech codec now superseded by Opus), and RealAudio, which made the first generation of internet radio possible over dial-up.

Why we care

A transcoder that only knows MP3 and AAC is not a transcoder; it is a vending machine. People bring us odd files: a broadcaster's BWF with metadata that must survive, a retro-game rip, a voicemail archive, a sound designer's Amiga sample. We would rather be the tool that can take all of them.

What chainDRiVE supports today (and what you can do with it)

Everything here is encoded and decoded on your device — in the browser or in the Android app — and every output can be previewed in the built-in player before you save it. Nothing is uploaded. For the formats we implemented ourselves (the broadcast containers, MPEG Layer I/II, and the console and telephony ADPCM family) we implemented the encoders and decoders ourselves and cross-checked them against ffmpeg to make sure we agree with the rest of the world.

  • Mainstream lossy: MP3 (LAME), AAC-LC, HE-AAC v1 and v2 (as .m4a or ADTS), Opus, Ogg Vorbis, WebM (Opus).
  • Lossless and PCM: FLAC, WAV PCM (integer and float), AIFF, AU, and headerless raw PCM.
  • Broadcast and containers: BWF with bext metadata, RF64 / BW64, Wave64, CAF, Creative VOC, Amiga 8SVX (including Fibonacci-delta).
  • MPEG Layer I and II: .mp1 and .mp2, plus MPEG audio inside WAV or BWF.
  • Lossless and audiophile family (WebAssembly): Monkey's Audio (.ape), WavPack (.wv), TTA (.tta) and Apple Lossless (ALAC in .m4a), for both reading and writing, plus Musepack (.mpc) and Speex (.spx). These load only when you use them.
  • AC-3 and ATRAC outputs: AC-3, E-AC-3 (Enhanced AC-3), ATRAC3 (LP2, LP4 and 105 kbps, as .at3 or .oma) and ATRAC1 (.aea, MiniDisc SP).
  • Open-only formats: DSD (DSF and DSDIFF, decimated to PCM), Shorten, TAK, WMA (including Pro, Lossless and Voice), RealAudio, DTS, ATRAC3+ and ATRAC9, Bink Audio and Smacker audio. There are no open encoders for these, so they are import only.
  • Game audio: BRSTM, HCA, FSB, Wwise WEM, AWC, CRI ADX and several hundred more extensions via vgmstream.
  • Tracker and DOS music: MOD, XM, S3M, IT and dozens of other tracker formats, plus MIDI, Doom MUS, XMI, IMF, CMF, DRO and RAD, rendered through an emulated OPL3 FM chip (the Sound Blaster and AdLib sound). These are note data, so we render them to audio and you can process and export the result.
  • Retro PC and console game audio (export): Sierra SOL, Westwood AUD, Maxis and EA XA, LucasArts VIMA, Doom DMX, plus console sample formats: SNES BRR, NES DMC, Neo Geo ADPCM-A and B, Mega Drive DAC and GBA Sappy.
  • Game archives and chiptunes (import): open PAK, PK3, PK4, WAD, VPK, BSA, BA2, MPQ, GRP, HOG, LAB, BIG and Sims .package archives, pick a sound and preview it before adding it to the queue. SPC, NSF, VGM, GBS and other chiptune dumps are rendered to audio, and a ROM scanner can find samples inside console ROM images you own. All of it runs on your device.
  • SBR Crystallizer (effect): an optional, off-by-default effect that rebuilds the high frequencies a narrow-band source or aggressive codec has lost, with controls for level, tilt, cutoff and more.
  • PlayStation and game ADPCM: XA, VAG / VB, Nintendo DSP-ADPCM, CRI ADX, and Yamaha ADPCM (raw and in WAV).
  • Voice and telephony: AMR-NB and AMR-WB, A-law and μ-law, GSM 6.10, IMA and Microsoft ADPCM, OKI/Dialogic VOX, G.726 and G.722 (G.726 is 8 kHz mono, G.722 is 16 kHz mono).

What can you actually do with that? Turn a folder of mixed recordings into one consistent format with a single loudness target; make a retro-sounding PlayStation or telephony version of a sound as a creative effect; produce a broadcast-ready BWF or MP2 for a station's playout system; open a game rip and convert it into something your phone can play; or just listen to how much a codec changes a sound by switching between them and comparing. The effects chain, LUFS normalisation and batch ZIP all work with any of them. There is a full summary on the chainDRiVE page, or try it in your browser.

What we want to try next

Most of the old wish list has landed. What is left is honest bench work, none of it a promise:

  • mp3PRO: files already open as plain MP3, but the SBR high-frequency layer was never documented, so playing it properly would be a reverse-engineering project.
  • OptimFROG, Oodle and RAD's newer audio codecs: no open decoder exists that we can ship.
  • ATRAC3plus and ATRAC9 export: The encoders for these are proprietary, so for now these are read-only.
  • MLP / TrueHD, the UT99 audio package format, and SID chiptune rips are the next candidates.

We would rather say so than put a button in the menu that does nothing. As always, we will tell you what is real, what is partial, and what is still on the bench.

Related: chainDRiVE's Format Support Grows Up, Codec Depth, chainDRiVE, and CHAiNAMP and Why Audio Transcoders Still Matter. The product: chainDRiVE (or open the web app).