Skip to content

[Porting Suggestion] Higgs-Audio-v3 TTS — looks like a solid candidate. #198

Description

@Rafa00127

Higgs-Audio-v3 TTS
The latest TTS model released by boson.ai.
It's a model across 100+ languages, with zero-shot voice cloning and emotion control. Here is a review video made by a friend from bilibili – the model shows not only high quality output, but also great emotion (and the video is quite funny because of it).
Benchmark results using a visual novel test set from the video:
Image

Anyway, it's probably the SOTA TTS model so far. The only problem is that the model is super huge (about 10GB). With a GGML implementation and quantization, it will be able to run on mid-range or even lower-end hardware. So I hope you can take a look at it when you have time, thanks!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions