Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

MIDI SVS mode may produce uneven rhythms

Đang mở
#60 1 bình luận 8 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Đình trệ
Công nghệ
python

Hướng nghiên cứu

Bắt đầu bằng cách lần theo logic suy luận độ dài âm vị MIDI SVS được mô tả trong issue và tái hiện nó bằng lời bài hát demo, chuỗi nốt và các khoảng thời lượng được cung cấp. So sánh nhịp điệu được tạo ra với âm thanh có thời lượng cố định được cung cấp và âm thanh X Studio; được xem là hoàn tất khi thời điểm bắt đầu của nguyên âm và thời lượng của phụ âm tạo ra timing nốt nhất quán trong chế độ MIDI SVS.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

enhancement must-read

Hello and thank you for your great work. However, I tried MIDI SVS of DiffSinger and found that there might be a conceptual mistake in the phoneme duration inference logic, which may lead to uneven rhythms of the output voice.
This possible mistake relates to the definitions of "note duration". Here I would like to show several examples.

Explanations of the duration of notes

As shown in the picture below, the duration of a note (containing one single syllable) is normally defined as the duration between the beginning of its vowel part and the beginning of the vowel part of the next note.
duration_of_notes

That is to say, notes begin at the beginning of their VOWEL parts, not their CONSONANT parts (as notes in MIDI SVS of DiffSinger currently do). When we sing, the rhythm sounds correct because every vowel starts on its right place, but not because consonants do; in fact, the length of consonants may affect the strength we feel, but theoretically not the rhythm.

Consequences of this kind of inconsistency

This kind of inconsistency can lead to chaotic rhythms. Take the demo lyric "小酒窝长睫毛是你最美的记号" for an example, and here is its music score:
music_score

Thus, we input:

input text
小 酒 窝 长 睫 毛 SP 是 你 最 美 的 记 号

input note
C#4 | F#4 | G#4 | A#4 F#4 | F#4 C#4 | C#4 | rest | C#4 | A#4 | G#4 | A#4 | G#4 | F#4 | C#4 

input duration
0.315789 | 0.315789 | 0.315789 | 0.315789 0.315789 | 0.315789 0.315789 | 0.315789 | 0.315789 | 0.315789 | 0.315789 | 0.315789 | 0.315789 | 0.315789 | 0.315789 | 0.315789 

The output audio sounds wired and is probably not sung in rhythm ("小酒窝_diffsinger_raw.wav" in the attachment).

I then used other algorithm to predict the duration of each phone, and tried to fix this incorrect rhythm:

input text
小 酒 窝 长 睫 毛 SP 是 你 最 美 的 记 号

input note
C#4 | F#4 | G#4 | A#4 F#4 | F#4 C#4 | C#4 | rest | C#4 | A#4 | G#4 | A#4 | G#4 | F#4 | C#4 

input duration
0.390789 | 0.375789 | 0.25579 | 0.420789 0.210789 | 0.420789 0.21079 | 0.420789 | 0.13579 | 0.405789 | 0.30079 | 0.330789 | 0.36079 | 0.25579 | 0.315789 | 0.42079 

The output audio sounds much better ("小酒窝_diffsinger_ fixed_phone_durations.wav" in the attachment).
However, as only the beginning of consonant parts, but not the vowel parts, can be specified in MIDI SVS mode of DiffSinger, we may never get correct rhythms (in theory).

As a comparison, I produced a piece of audio with X Studio (Xiaoice Sing) that has the correct rhythm ("小酒窝_xiaoicesing_correct_rhythm.wav" in the attachment).

Here are the audios: audios.zip

My expectations

My teammates and I are trying to bring DiffSinger to more ordinary fans and users of SVS technology and products. These people (or you can say, most people) are more familiar with the interaction mode that takes notes or music scores as input. Therefore, correct rhythms are important and can help a lot.
It helps a lot if you fix the issue in rhythms (i. e. specify the beginning of vowels and predict the duration of the consonants).
I'm looking forward to your improvements.

Ngôn ngữ chính
Python
Star
4.9k
Fork
826
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của MoonInTheRiver/DiffSinger

Tất cả issue của MoonInTheRiver/DiffSinger

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.