Audio Is Generated, Not Added
Dialogue, ambience and effects come out of the same pass as the picture, so they match what is on screen. Every other model here returns a silent clip you then have to score.
Veo 3.1 is Google DeepMind's video model, and the only one in this catalog that treats audio as part of the generation rather than something you add afterwards. It also takes reference images for subject consistency, which is what lets it sit inside a reel instead of only producing standalone clips.
Mô hìnhDialogue, ambience and effects come out of the same pass as the picture, so they match what is on screen. Every other model here returns a silent clip you then have to score.
Veo 3.1 accepts reference images for consistent subject appearance, which is why it qualifies for the reel pipeline. A model that only reads a start frame cannot hold a face past shot one.
Up to 4K at 16:9 or 9:16, 4 to 8 seconds. The most expensive model in the catalog and the one to reach for when the shot has to carry the reel.
Đăng nhập là đã có sẵn 100 tín dụng đầu tiên — không cần thẻ.