Those are the only valid values.
Case matters. The model interpolates whatever it is given straight into
its prompt, so ananya is not the same request as Ananya — it produces
audibly different audio. Anything outside the two above is rejected.
How much audio a clip is
Roughly 800 characters per minute, measured across both voices. So 10,000
characters is about 12.5 minutes.