Computers break examples of human speech into small bits of sound, then put them together to sound like people do.