深度学习自动吉他转录(使用卷积神经网络加快音乐学习)
loadstar_cn
编辑于 2022年01月28日 13:13
收录于文集
共2篇

这篇文章概述了使用Python、TensorFlow和Keras从音频文件自动转录吉他的实现,并详细介绍了执行的表面级方法。为了进行学习,GuitarSet数据集(https://zenodo.org/record/1422265#.XQvsmohKi01)使用了大量独立的吉他录音,带有相应的标签。请注意,本项目的大部分方向都是由NEMISIG 2019的研究海报(http://nemisig2019.nemisig.org/images/kimSlides.pdf)提供的。

照片来源: Jacek Dylag on Unsplash

背景

如果你熟悉卷积神经网络(CNN)(https://towardsdatascience.com/a-comprehensive-guide-to-convolutional-neural-networks-the-eli5-way-3bd2b1164a53),那么你可能听说过它们在用于计算机视觉的图像处理和分析方面的潜力。 我们想要利用 CNN 的这种功能来输出吉他谱; 因此,首先需要使用 onstant-Q 变换将输入音频文件转换为频谱图图像。

为什么是常数Q变换?

为了理解使用恒Q变换而不是傅里叶变换来选择频率和创建输入图像的好处,我们必须检查音符是如何定义的:

图片来源:https://www.soundonsound.com/forum/viewtopic.php?f=16&t=64005

在“音符”列中,音符由左侧的字母标识,右侧的数字表示当前的八度。下面是音符C前六个八度的频率(Hz)图。

我们可以看到,每个连续的八度音阶是前一个八度音阶频率的两倍。由于一个八度音程跨越十二个音符,我们知道频率必须每十二个音符加倍,这可以用下面的公式[1]表示:

通过绘制这种关系,我们可以看到下图显示了一条指数曲线:

由于这种指数性质,常数Q变换比傅里叶变换更适合拟合音乐数据,因为它的输出是振幅与对数频率的关系。此外,恒Q变换的精度类似于对数尺度,并模仿人耳,在较低频率下具有较高频率分辨率,在较高频率下具有较低的分辨率[1]。

在Python中应用常量Q变换

使用libROSA库,可以轻松地将常量-Q转换应用于python中的音频文件。

代码块
Python
自动换行
复制代码
def audio_CQT(file_num, start, dur):  # start and dur in seconds
    
    # Load audio and define paths
    path = r'C:.../GuitarSet_annotation_and_audio/GuitarSet/audio/audio_hex-pickup_debleeded/'
    audio_file = os.listdir(path)
    audio_path = os.path.join(path, audio_file[file_num])

    # Function for removing noise
     def cqt_lim(CQT):
        new_CQT = np.copy(CQT)
        new_CQT[new_CQT < -60] = -120
        return new_CQT
    
    # Perform the Constant-Q Transform
    data, sr = librosa.load(audio_path, sr = None, mono = True, offset = start, duration = dur)
    CQT = librosa.cqt(data, sr = 44100, hop_length = 1024, fmin = None, n_bins = 96, bins_per_octave = 12)
    CQT_mag = librosa.magphase(CQT)[0]**4
    CQTdB = librosa.core.amplitude_to_db(CQT_mag, ref = np.amax)
    new_CQT = cqt_lim(CQTdB)
复制成功

通过从指定的开始时间 (start) 到持续时间 (dur) 遍历 GuitarSet 数据集中的每个音频文件并将输出保存为图像,我们可以创建训练 CNN 所需的输入图像。 对于这个项目,dur 置为 0.2 秒,start 设置为从零增加到每个音频文件的长度,按设置的持续时间,这可以产生以下结果:

请注意,当用作CNN的输入图像时,颜色方案首先转换为灰度。

训练方案

对于每个恒定Q变换图像,必须有一个方法以便网络可以调整其猜测。幸运的是,GuitarSet数据集包含作为MIDI值播放的所有音符、每个音符开始录制的时间,以及每个音频文件的音持续时间。注意:下面的代码片段放在一个函数中,以便每0.2秒音频就可以使用它们。

首先,必须从jams文件中提取在加载的0.2秒音频中播放的唯一音符(作为MIDI音符检索)。

代码块
Python
自动换行
复制代码
# Initialize variables
cnt_row = -1
cnt_col = 0
cnt_zero = 0

# Grab all relevant MIDI data (available in MIDI_dat)
for i in range(0, len(jam['annotations'])):
    if jam['annotations'][int(i)]['namespace'] == 'note_midi':
        for j in range(0, len(sorted(jam['annotations'][int(i)]['data']))):
            cnt_row = cnt_row + 1
            for k in range(0, len(sorted(jam['annotations'][int(i)]['data'])[int(j)]) - 1):
                if cnt_zero == 0:
                    MIDI_arr = np.zeros((len(sorted(jam['annotations'][int(i)]['data'])), len(sorted(jam['annotations'][int(i)]['data'])[int(j)]) - 1), dtype = np.float32)
                    cnt_zero = cnt_zero + 1
                if cnt_zero > 0:
                    MIDI_arr = np.vstack((MIDI_arr, np.zeros((len(sorted(jam['annotations'][int(i)]['data'])), len(sorted(jam['annotations'][int(i)]['data'])[int(j)]) - 1), dtype = np.float32)))
                    cnt_zero = cnt_zero + 1  # Keep
                if cnt_col > 2:
                    cnt_col = 0
                MIDI_arr[cnt_row, cnt_col] = sorted(jam['annotations'][int(i)]['data'])[int(j)][int(k)]
                cnt_col = cnt_col + 1
MIDI_dat = np.zeros((cnt_row + 1, cnt_col), dtype = np.float32)
cnt_col2 = 0
for n in range(0, cnt_row + 1):
    for m in range(0, cnt_col):
        if cnt_col2 > 2:
            cnt_col2 = 0
        MIDI_dat[n, cnt_col2] = MIDI_arr[n, cnt_col2]
        cnt_col2 = cnt_col2 + 1
        
 # Return the unique MIDI notes played (available in MIDI_val)
MIDI_dat_dur = np.copy(MIDI_dat)
for r in range(0, len(MIDI_dat[:, 0])):
    MIDI_dat_dur[r, 0] = MIDI_dat[r, 0] + MIDI_dat[r, 1]
tab_1, = np.where(np.logical_and(MIDI_dat[:, 0] >= start, MIDI_dat[:, 0] <= stop))
tab_2, = np.where(np.logical_and(MIDI_dat_dur[:, 0] >= start, MIDI_dat_dur[:, 0] <= stop))
tab_3, = np.where(np.logical_and(np.logical_and(MIDI_dat[:, 0] < start, MIDI_dat_dur[:, 0] > stop), MIDI_dat[:, 1] > int(stop-start)))
if tab_1.size != 0 and tab_2.size == 0 and tab_3.size == 0:
    tab_ind = tab_1
if tab_1.size == 0 and tab_2.size != 0 and tab_3.size == 0:
    tab_ind = tab_2
if tab_1.size == 0 and tab_2.size == 0 and tab_3.size != 0:
        tab_ind = tab_3
if tab_1.size != 0 and tab_2.size != 0 and tab_3.size == 0:
    tab_ind = np.concatenate([tab_1, tab_2])
if tab_1.size != 0 and tab_2.size == 0 and tab_3.size != 0:
    tab_ind = np.concatenate([tab_1, tab_3])
if tab_1.size == 0 and tab_2.size != 0 and tab_3.size != 0:
    tab_ind = np.concatenate([tab_2, tab_3])
if tab_1.size != 0 and tab_2.size != 0 and tab_3.size != 0:
    tab_ind = np.concatenate([tab_1, tab_2, tab_3])
if tab_1.size == 0 and tab_2.size == 0 and tab_3.size == 0:
    tab_ind = []
if len(tab_ind) != 0:
    MIDI_val = np.zeros((len(tab_ind), 1), dtype = np.float32)
    for z in range(0, len(tab_ind)):
        MIDI_val[z, 0] = int(round(MIDI_dat[tab_ind[z], 2]))
elif len(tab_ind) == 0:
    MIDI_val = []
MIDI_val = np.unique(MIDI_val)
if MIDI_val.size >= 6:
    MIDI_val = np.delete(MIDI_val, np.s_[6::])
复制成功

一次只能播放六个可能的音符(每根弦最多一个音符);因此,代码通常会重复六次。

代码块
Python
自动换行
复制代码
# Initialize variables
Fret = np.zeros((6, 18), dtype = np.int32)
Sol = np.copy(Fret)
fcnt = -1
fcnt2 = 0

# Retrieve all possible notes played
for q in range(0, 6):
    for e in range(0, 18):
        if q == 0:
            Fret[q, e] = 40 + e
        elif q == 1:
            Fret[q, e] = 45 + e
        elif q == 2:
            Fret[q, e] = 50 + e
        elif q == 3:
            Fret[q, e] = 55 + e
        elif q == 4:
            Fret[q, e] = 59 + e
        elif q == 5:
            Fret[q, e] = 64 + e
for t in range(0, len(MIDI_val)):
    Fret_played = (Fret == int(MIDI_val[t]))
    fcnt = fcnt + 1
    cng = 0
    for dr in range(0, len(Fret[:, 0])):
        for dc in range(0, len(Fret[0, :])):
            if Fret_played[dr, dc]*1 == 1:
                if cng == 0:
                    fcnt2 = 0
                    cng = cng + 1
                f_row[fcnt, fcnt2] = dr
                f_col[fcnt, fcnt2] = dc
                fcnt2 = fcnt2 + 1
            Fret_played[dr, dc] = Fret_played[dr, dc]*1
            if Fret_played[dr, dc] == 1:
                Sol[dr, dc] = Fret_played[dr, dc]
复制成功

首先,在变量Fret下创建一个MIDI值矩阵(6,18),该矩阵表示吉他的六根弦和18个Fret:

代码块
Python
自动换行
复制代码
[[0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0]
 [0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 1]
 [0 1 0 0 0 0 0 0 0 0 0 0 1 0 0 1 0 0]
 [0 0 0 0 0 0 0 1 0 0 1 0 0 0 0 0 0 0]
 [0 0 0 1 0 0 1 0 0 0 0 0 0 0 0 0 0 0]
 [0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]]
复制成功

然后使用 Fret 确定在吉他上检索到的唯一音符的所有可能位置,下面的矩阵显示了一个可能的解决方案:

代码块
Python
自动换行
复制代码
[[0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0]
 [0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 1]
 [0 1 0 0 0 0 0 0 0 0 0 0 1 0 0 1 0 0]
 [0 0 0 0 0 0 0 1 0 0 1 0 0 0 0 0 0 0]
 [0 0 0 1 0 0 1 0 0 0 0 0 0 0 0 0 0 0]
 [0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]]
复制成功

必须确定FRET和琴弦组合的所有可能解决方案。 创造了“手指经济”的概念——和弦的最低音符,即根音,与和弦中的其余音符进行比较,其中每个音符来自根音的音品数量(不考虑弦)是 加起来创造了一个“手指经济”的数字。 选择具有最低“手指经济”数的解决方案作为正确的和弦形状。

代码块
Python
自动换行
复制代码
# Initialize variables
f_row = np.full((6, 6), np.inf)  # 6 strings with 1 note per string
f_col = np.full((6, 6), np.inf)

# Initialize the 6 possible note solutions (one note per string)
f_sol_0 = np.copy(f_col)
f_sol_1 = np.copy(f_col)
f_sol_2 = np.copy(f_col)
f_sol_3 = np.copy(f_col)
f_sol_4 = np.copy(f_col)
f_sol_5 = np.copy(f_col)
pri_cnt_c, = np.where(np.isfinite(f_col[0, :]))
pri_cnt_r, = np.where(np.isfinite(f_col[:, 0]))
if len(MIDI_val) > 1:
    for pri in range(0, len(pri_cnt_c)):
        for sub_r in range(1, 6):
            for sub_c in range(0, len(f_sol_0[0, :])):
                if pri == 0:
                    f_sol_0[sub_r, sub_c] = abs(f_col[0, pri] - f_col[sub_r, sub_c])
                if pri == 1:
                    f_sol_1[sub_r, sub_c] = abs(f_col[0, pri] - f_col[sub_r, sub_c])
                if pri == 2:
                    f_sol_2[sub_r, sub_c] = abs(f_col[0, pri] - f_col[sub_r, sub_c])
                if pri == 3:
                    f_sol_3[sub_r, sub_c] = abs(f_col[0, pri] - f_col[sub_r, sub_c])
                if pri == 4:
                    f_sol_4[sub_r, sub_c] = abs(f_col[0, pri] - f_col[sub_r, sub_c])
                if pri == 5:
                    f_sol_5[sub_r, sub_c] = abs(f_col[0, pri] - f_col[sub_r, sub_c])
if len(pri_cnt_r) == 0 or len(pri_cnt_c) == 0:
    True_tab = np.copy(np.zeros((6, 18), dtype = np.int32))
else:
    ck_sol_0 = np.zeros((len(pri_cnt_r) - 1, len(pri_cnt_c) - 1), dtype = np.int32)
    sol_ind_0 = np.copy(ck_sol_0)
    ck_sol_1 = np.zeros((len(pri_cnt_r) - 1, len(pri_cnt_c) - 1), dtype = np.int32)
    sol_ind_1 = np.copy(ck_sol_0)
    ck_sol_2 = np.zeros((len(pri_cnt_r) - 1, len(pri_cnt_c) - 1), dtype = np.int32)
    sol_ind_2 = np.copy(ck_sol_0)
    ck_sol_3 = np.zeros((len(pri_cnt_r) - 1, len(pri_cnt_c) - 1), dtype = np.int32)
    sol_ind_3 = np.copy(ck_sol_0)
    ck_sol_4 = np.zeros((len(pri_cnt_r) - 1, len(pri_cnt_c) - 1), dtype = np.int32)
    sol_ind_4 = np.copy(ck_sol_0)
    ck_sol_5 = np.zeros((len(pri_cnt_r) - 1, len(pri_cnt_c) - 1), dtype = np.int32)
    sol_ind_5 = np.copy(ck_sol_0)

    # Replace infinite values with high finite values for each solution
    for ck_sol in range(0, len(pri_cnt_c)):
        for pri_sol_r in range(1, len(pri_cnt_r)):
            for pri_sol_c in range(0, len(pri_cnt_c) - 1):  # Random - 1
                if ck_sol == 0:
                    if np.any(np.isinf(f_sol_0[pri_sol_r, :])):
                        avoid_0 = np.argwhere(np.isinf(f_sol_0[pri_sol_r, :]))
                        f_sol_0[pri_sol_r, avoid_0] = 999
                if ck_sol == 1:
                    if np.any(np.isinf(f_sol_1[pri_sol_r, :])):
                        avoid_1 = np.argwhere(np.isinf(f_sol_1[pri_sol_r, :]))
                        f_sol_1[pri_sol_r, avoid_1] = 999
                    ck_sol_1[0, pri_sol_c] = min(f_sol_1[pri_sol_r, :])
                if ck_sol == 2:
                    if np.any(np.isinf(f_sol_2[pri_sol_r, :])):
                        avoid_2 = np.argwhere(np.isinf(f_sol_2[pri_sol_r, :]))
                        f_sol_2[pri_sol_r, avoid_2] = 999
                    ck_sol_2[0, pri_sol_c] = min(f_sol_2[pri_sol_r, :])
                if ck_sol == 3:
                    if np.any(np.isinf(f_sol_3[pri_sol_r, :])):
                        avoid_3 = np.argwhere(np.isinf(f_sol_3[pri_sol_r, :]))
                        f_sol_3[pri_sol_r, avoid_3] = 999
                    ck_sol_3[0, pri_sol_c] = min(f_sol_3[pri_sol_r, :])
                if ck_sol == 4:
                    if np.any(np.isinf(f_sol_4[pri_sol_r, :])):
                        avoid_4 = np.argwhere(np.isinf(f_sol_4[pri_sol_r, :]))
                        f_sol_4[pri_sol_r, avoid_4] = 999
                    ck_sol_4[0, pri_sol_c] = min(f_sol_4[pri_sol_r, :])
                if ck_sol == 5:
                    if np.any(np.isinf(f_sol_5[pri_sol_r, :])):
                        avoid_5 = np.argwhere(np.isinf(f_sol_5[pri_sol_r, :]))
                        f_sol_5[pri_sol_r, avoid_5] = 999
                    ck_sol_5[0, pri_sol_c] = min(f_sol_5[pri_sol_r, :])

    # Determine "rating" for each solution
    tab_sol_0 = np.argmin(f_sol_0, axis = 1)
    min_sol_0 = np.min(f_sol_0, axis = 1)
    if np.any(np.isinf(min_sol_0[:])):
        rep_0 = np.argwhere(np.isinf(min_sol_0[:]))
        min_sol_0[rep_0] = 0
    tab_sol_1 = np.argmin(f_sol_1, axis = 1)
    min_sol_1 = np.min(f_sol_1, axis = 1)
    if np.any(np.isinf(min_sol_1[:])):
        rep_1 = np.argwhere(np.isinf(min_sol_1[:]))
        min_sol_1[rep_1] = 0
    tab_sol_2 = np.argmin(f_sol_2, axis = 1)
    min_sol_2 = np.min(f_sol_2, axis = 1)
    if np.any(np.isinf(min_sol_2[:])):
        rep_2 = np.argwhere(np.isinf(min_sol_2[:]))
        min_sol_2[rep_2] = 0
    tab_sol_3 = np.argmin(f_sol_3, axis = 1)
    min_sol_3 = np.min(f_sol_3, axis = 1)
    if np.any(np.isinf(min_sol_3[:])):
        rep_3 = np.argwhere(np.isinf(min_sol_3[:]))
        min_sol_3[rep_3] = 0
    tab_sol_4 = np.argmin(f_sol_4, axis = 1)
    min_sol_4 = np.min(f_sol_4, axis = 1)
    if np.any(np.isinf(min_sol_4[:])):
        rep_4 = np.argwhere(np.isinf(min_sol_4[:]))
        min_sol_4[rep_4] = 0
    tab_sol_5 = np.argmin(f_sol_5, axis = 1)
    min_sol_5 = np.min(f_sol_5, axis = 1)
    if np.any(np.isinf(min_sol_5[:])):
        rep_5 = np.argwhere(np.isinf(min_sol_5[:]))
        min_sol_5[rep_5] = 0
    sol_0 = np.sum(min_sol_0[:])
    sol_1 = np.sum(min_sol_1[:])
    sol_2 = np.sum(min_sol_2[:])
    sol_3 = np.sum(min_sol_3[:])
    sol_4 = np.sum(min_sol_4[:])
    sol_5 = np.sum(min_sol_4[:])
复制成功

虽然这种方法并不总是与录音中演奏的和弦的正确版本相匹配,但它不会对 CNN 的性能产生负面影响,因为在开放位置演奏的 C 大调和弦与在 8 品上演奏的 C 大调没有区别。 随后,使用最终解决方案中的弦和 frets数组组合选择最终解决方案:

代码块
Python
自动换行
复制代码
# Initalize variables
acc_sol = False
idx_pass = False

# Choose best solution based on previous rating
if len(pri_cnt_c) == 1:
    fin_sol_arr = sol_0
if len(pri_cnt_c) == 2:
    fin_sol_arr = np.append(sol_0, sol_1)
if len(pri_cnt_c) == 3:
    fin_sol_arr = np.append(np.append(sol_0, sol_1), sol_2)
if len(pri_cnt_c) == 4:
    fin_sol_arr = np.append(np.append(sol_0, sol_1), np.append(sol_2, sol_3))
if len(pri_cnt_c) == 5:
    fin_sol_arr = np.array(np.append(np.append(sol_0, sol_1), np.append(sol_2, sol_3)), sol_4)
if len(pri_cnt_c) == 6:
    fin_sol_arr = np.array(np.append(np.append(sol_0, sol_1), np.append(sol_2, sol_3)), np.append(sol_4, sol_5))
fin_choice = np.argmin(fin_sol_arr)
response, ret_cnts, ret_idx = np.unique(fin_sol_arr, return_counts = True, return_index = True)
ret_idx = [np.argwhere(idx_cnt == fin_sol_arr) for idx_cnt in np.unique(fin_sol_arr)]
for idx_cnt_row in range(0, len(ret_idx)):
    if np.amin(response) == np.amin(fin_sol_arr) and len(ret_idx[idx_cnt_row]) > 2:
        fin_sol_arr = np.delete(fin_sol_arr, np.argwhere(np.amin(fin_sol_arr)))
if np.amin(response) == np.amin(fin_sol_arr) and ret_cnts[np.argwhere(np.amin(fin_sol_arr))] > 2:
    fin_sol_arr = np.delete(fin_sol_arr, np.argwhere(np.amin(fin_sol_arr)))
    fin_choice = np.argmin(fin_sol_arr)

# Choose solution and choose the next best solution if there are two notes on one string
while acc_sol == False:
    fin_tab_row = np.zeros((len(pri_cnt_r)), dtype = np.int32)
    fin_tab_col = np.zeros((len(pri_cnt_r)), dtype = np.int32)
    if fin_choice == 0:
        fin_tab_row[0] = f_row[0, 0]
        fin_tab_col[0] = f_col[0, 0]
        for counter in range(1, len(pri_cnt_r)):
            fin_tab_row[counter] = f_row[counter, tab_sol_0[counter]]
            fin_tab_col[counter] = f_col[counter, tab_sol_0[counter]]
    if fin_choice == 1:
        fin_tab_row[0] = f_row[0, 1]
        fin_tab_col[0] = f_col[0, 1]
        for counter in range(1, len(pri_cnt_r)):
            fin_tab_row[counter] = f_row[counter, tab_sol_1[counter]]
            fin_tab_col[counter] = f_col[counter, tab_sol_1[counter]]
    if fin_choice == 2:
        fin_tab_row[0] = f_row[0, 2]
        fin_tab_col[0] = f_col[0, 2]
        for counter in range(1, len(pri_cnt_r)):
            fin_tab_row[counter] = f_row[counter, tab_sol_2[counter]]
            fin_tab_col[counter] = f_col[counter, tab_sol_2[counter]]
    if fin_choice == 3:
        fin_tab_row[0] = f_row[0, 3]
        fin_tab_col[0] = f_col[0, 3]
        for counter in range(1, len(pri_cnt_r)):
            fin_tab_row[counter] = f_row[counter, tab_sol_3[counter]]
            fin_tab_col[counter] = f_col[counter, tab_sol_3[counter]]
    if fin_choice == 4:
        fin_tab_row[0] = f_row[0, 4]
        fin_tab_col[0] = f_col[0, 4]
        for counter in range(1, len(pri_cnt_r)):
            fin_tab_row[counter] = f_row[counter, tab_sol_4[counter]]
            fin_tab_col[counter] = f_col[counter, tab_sol_4[counter]]
    if fin_choice == 5:
        fin_tab_row[0] = f_row[0, 5]
        fin_tab_col[0] = f_col[0, 5]
        for counter in range(1, len(pri_cnt_r)):
            fin_tab_row[counter] = f_row[counter, tab_sol_5[counter]]
            fin_tab_col[counter] = f_col[counter, tab_sol_5[counter]]
    acc_sol = True
    idx_cnt = [np.argwhere(uni_cnt == fin_tab_row) for uni_cnt in np.unique(fin_tab_row)]
    max_len_cnt = np.zeros((len(idx_cnt)), dtype = np.int32)
    for str_cnt_row in range(0, len(idx_cnt)):
        if len(idx_cnt[str_cnt_row]) > 1:
            fin_sol_arr = fin_sol_arr.astype('int64')
            if fin_sol_arr.size > 1:
                fin_sol_arr = np.delete(fin_sol_arr, fin_choice)
                idx_pass = True
                acc_sol = False
                break
            else:
                continue
    fin_choice = np.argmin(fin_sol_arr)
fin_tab_row = abs(fin_tab_row - 5)

# Return the final tab
True_tab = np.copy(np.zeros((6, 18), dtype = np.int32))
for tt_cnt in range(0, len(fin_tab_col)):
    True_tab[fin_tab_row[tt_cnt], fin_tab_col[tt_cnt]] = 1
复制成功

此外,在每行的第一列中,如果存在注释(行中为 1),则附加零,如果不存在注释,反之亦然。 这样做是为了让 softmax 函数仍然可以在没有播放音符的情况下为字符串选择一个别。

前面的代码片段返回数据,使得输出类似于类别的 one-hot 编码,每 0.2 秒的音频返回以下矩阵格式:

代码块
Python
自动换行
复制代码
[[1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]
 [1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]
 [0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0]
 [0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0]
 [0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0]
 [0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0]]
复制成功

上面的矩阵是从 GuitarSet 数据集中随机选择一个 0.2 秒的吉他谱解决方案。 每个矩阵形状是 (6, 19),其中六行对应于每个吉他弦(从上到下的 eBGDAE)。 第一列标识该弦是未在播放,第二列标识是否正在播放开放字符串,第三至第十九列标识从第一个音品开始正在播放的特定音品。 训练时,这个矩阵被分解成六个独立的数组来训练模型的每个头部。

模型架构和训练

Keras 功能 API 用于创建以下多任务分类模型,其中训练和测试数据的比例为 90/10。 该模型具有六个任务(eBGDAE 弦乐),以确定弦乐是否未播放、打开或正在播放音符。 对于六个输出中的每一个,都应用了 softmax 激活和分类交叉熵损失函数。 添加了 Dropout 层以减少过度拟合。

代码块
Python
自动换行
复制代码
# Optimizer
epochs = 30
learning_rate = 0.01 
momentum = 0.8
decay = learning_rate/epochs
sgd = SGD(lr = learning_rate, momentum = momentum, decay = decay, nesterov = False)

# Training (Functional Method)
model_in = Input(shape = input_shape)
conv1 = Conv2D(32, kernel_size = (3, 3), activation = 'relu')(model_in)
conv2 = Conv2D(64, kernel_size = (3, 3), activation = 'relu')(conv1)
conv3 = Conv2D(64, kernel_size = (3, 3), activation = 'relu')(conv2)
pool1 = MaxPooling2D(pool_size = (2, 2), strides = (2, 2))(conv3)
flat = Flatten()(pool1)

# Create fully connected model heads
y1 = Dense(152, activation = 'relu')(flat)
y1 = Dropout(0.5)(y1)
y1 = Dense(76)(y1)
y1 = Dropout(0.2)(y1)

y2 = Dense(152, activation = 'relu')(flat)
y2 = Dropout(0.5)(y2)
y2 = Dense(76)(y2)
y2 = Dropout(0.2)(y2)

y3 = Dense(152, activation = 'relu')(flat)
y3 = Dropout(0.5)(y3)
y3 = Dense(76)(y3)
y3 = Dropout(0.2)(y3)

y4 = Dense(152, activation = 'relu')(flat)
y4 = Dropout(0.5)(y4)
y4 = Dense(76)(y4)
y4 = Dropout(0.2)(y4)

y5 = Dense(152, activation = 'relu')(flat)
y5 = Dropout(0.5)(y5)
y5 = Dense(76)(y5)
y5 = Dropout(0.2)(y5)

y6 = Dense(152, activation = 'relu')(flat)
y6 = Dropout(0.5)(y6)
y6 = Dense(76)(y6)
y6 = Dropout(0.2)(y6)

# Connect heads to final output layer
out1 = Dense(19, activation = 'softmax', name = 'estring')(y1)
out2 = Dense(19, activation = 'softmax', name = 'Bstring')(y2)
out3 = Dense(19, activation = 'softmax', name = 'Gstring')(y3)
out4 = Dense(19, activation = 'softmax', name = 'Dstring')(y4)
out5 = Dense(19, activation = 'softmax', name = 'Astring')(y5)
out6 = Dense(19, activation = 'softmax', name = 'Estring')(y6)

# Create model
model = Model(inputs = model_in, outputs = [out1, out2, out3, out4, out5, out6])
model.compile(optimizer = sgd, loss = ['categorical_crossentropy','categorical_crossentropy','categorical_crossentropy', 'categorical_crossentropy','categorical_crossentropy','categorical_crossentropy'], metrics = ['accuracy'])
复制成功

该模型运行了30个epoch,并记录了每个字符串的最终精度。

代码块
Python
自动换行
复制代码
history = model.fit(x_train, [e_train, B_train, G_train, D_train, A_train, E_train], batch_size = batch_size, epochs = epochs, verbose = 1, validation_data = (x_test, [e_test, B_test, G_test, D_test, A_test, E_test]))

score = model.evaluate(x_test, [e_test, B_test, G_test, D_test, A_test, E_test], verbose = 1)
复制成功

结果

该模型没有使用整个 GuitarSet 数据集音频文件,但是使用了足够数量的输入文件,总共 40828 个训练图像和 4537 个测试样本。 每个弦的准确度被确定为:

代码块
Python
自动换行
复制代码
Test accuracy estring: 0.9336566011601407
Test accuracy Bstring: 0.8521049151158895
Test accuracy Gstring: 0.8283006392545786
Test accuracy Dstring: 0.7831165967256665
Test accuracy Astring: 0.8053780030331896
Test accuracy Estring: 0.8514436851100615
复制成功

结果平均准确率为 84.23%。

结论

这个模型还没有准备好开始创建全长吉他指法谱,因为一些问题仍然存在。 当前模型还没有考虑到一个音符的持续时间,并且会在代码中指定的持续时间内继续重复该选项卡。 此外,于和弦可以有不同的变体,每个变体都包含相同的音符,因此该模型无法识别何时使用特定的发声——这可能会带来不便——但这不是一个大问题。 然而,该模型能够正确标记音频片段能力是一项了不起的发展。

参考

[1] C. Schörkhuber and Anssi Klapuri, Constant-Q transform toolbox for music processing(https://www.researchgate.net/publication/228523955_Constant-Q_transform_toolbox_for_music_processing) (2010), 7th Sound and Music Computing Conference.

原文地址: https://towardsdatascience.com/audio-to-guitar-tab-with-deep-learning-d76e12717f81

作者:Darren Tio(https://medium.com/@darrentio)

翻译:网页链接​ 2022-01-28