简体   繁体   English

为什么 epoch 2 花费的时间是 epoch 1 的 18 倍?

[英]Why does epoch 2 take 18 times more time than epoch 1?

I have the following neural network in keras (the review of it is probably not necessary to answer my question:我在 keras 中有以下神经网络(对它的审查可能没有必要回答我的问题:

Short summary: It's a neural network that takes images as input and outputs images too.简短摘要:这是一个神经网络,将图像作为输入并输出图像。 The neural network is mostly convolutional.神经网络主要是卷积的。 I use generators.我使用发电机。 Also, I have two callbacks: one for TensorBoard and one for chechpoint saving另外,我有两个回调:一个用于 TensorBoard,另一个用于保存检查点

class modelsClass(object):
    def __init__(self, img_rows = 272, img_cols = 480):

        self.img_rows = img_rows
        self.img_cols = img_cols

    def addPadding(self, layer, level): #height, width, level):

        w1, h1 = self.img_cols, self.img_rows
        w2, h2 = int(w1/2), int(h1/2)
        w3, h3 = int(w2/2), int(h2/2)
        w4, h4 = int(w3/2), int(h3/2)
        h = [h1, h2, h3, h4]
        w = [w1, w2, w3, w4]

        # Target width and height
        tw = w[level-1]
        th = h[level-1]

        # Source width and height
        lsize = keras.int_shape(layer)
        sh = lsize[1]
        sw = lsize[2]

        pw = (0, tw - sw)
        ph = (0, th - sh)

        layer = ZeroPadding2D(padding=(ph, pw), data_format="channels_last")(layer)

        return layer

[I need to break the code with some text here to post the question] [我需要在这里用一些文字打破代码来发布问题]

    def getmodel(self):

        input_blurred = Input((self.img_rows, self.img_cols,3))

        conv1 = Conv2D(64, (3, 3), activation='relu', padding='same')(input_blurred)
        conv1 = Conv2D(64, (3, 3), activation='relu', padding='same')(conv1)
        pool1 = MaxPooling2D(pool_size=(2, 2))(conv1)

        conv2 = Conv2D(128, (3, 3), activation='relu', padding='same')(pool1)
        conv2 = Conv2D(128, (3, 3), activation='relu', padding='same')(conv2)
        pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)

        conv3 = Conv2D(256, (3, 3), activation='relu', padding='same')(pool2)
        conv3 = Conv2D(256, (3, 3), activation='relu', padding='same')(conv3)
        pool3 = MaxPooling2D(pool_size=(2, 2))(conv3)

        conv4 = Conv2D(512, (3, 3), activation='relu', padding='same')(pool3)
        conv4 = Conv2D(512, (3, 3), activation='relu', padding='same')(conv4)
        pool4 = MaxPooling2D(pool_size=(2, 2))(conv4)

        conv5 = Conv2D(1024, (3, 3), activation='relu', padding='same')(pool4)
        conv5 = Conv2D(1024, (3, 3), activation='relu', padding='same')(conv5)

        up6 = Conv2DTranspose(512, (2, 2), strides=(2, 2), padding='same')(conv5)
        up6 = self.addPadding(up6,level=4)
        up6 = concatenate([up6,conv4], axis=3)
        conv6 = Conv2D(512, (3, 3), activation='relu', padding='same')(up6)
        conv6 = Conv2D(512, (3, 3), activation='relu', padding='same')(conv6)

        up7 = Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same')(conv6)
        up7 = self.addPadding(up7,level=3)
        up7 = concatenate([up7,conv3], axis=3)
        conv7 = Conv2D(256, (3, 3), activation='relu', padding='same')(up7)
        conv7 = Conv2D(256, (3, 3), activation='relu', padding='same')(conv7)

        up8 = Conv2DTranspose(128, (2, 2), strides=(2, 2), padding='same')(conv7)
        up8 = self.addPadding(up8,level=2)
        up8 = concatenate([up8,conv2], axis=3)
        conv8 = Conv2D(128, (3, 3), activation='relu', padding='same')(up8)
        conv8 = Conv2D(128, (3, 3), activation='relu', padding='same')(conv8)

        up9 = Conv2DTranspose(64, (2, 2), strides=(2, 2), padding='same')(conv8)
        up9 = self.addPadding(up9,level=1)
        up9 = concatenate([up9,conv1], axis=3)
        conv9 = Conv2D(64, (3, 3), activation='relu', padding='same')(up9)
        conv9 = Conv2D(64, (3, 3), activation='relu', padding='same')(conv9)

        conv10 = Conv2D(3, (1, 1), activation='linear')(conv9)

        model = Model(inputs=input_blurred, outputs=conv10)

        return model

Then the code is:然后代码是:

models = modelsClass(720, 1280)
model = models.getmodel()

model.compile(optimizer='adam', loss='mean_absolute_error')
model_checkpoint = ModelCheckpoint('checkpoints/cp.ckpt', monitor='val_loss', verbose=0, save_best_only=False, save_weights_only=False, mode='auto', save_freq='epoch')
tensorboard_callback = tf.keras.callbacks.TensorBoard(log_dir='some_dir', histogram_freq=1)
model_history = model.fit_generator(generator_train, epochs=3,
                          steps_per_epoch=900,
                          callbacks=[tensorboard_callback, model_checkpoint],
                          validation_data=generator_val, validation_steps=100)

where generator_train.__len__ = 900 , generator_val.__len__ = 100 , batch size for both = 1.其中generator_train.__len__ = 900generator_val.__len__ = 100 ,两者的批量大小 = 1。
Time for epoch 1 is 10 minutes, while epoch 2 takes 3 hours. epoch 1 的时间为 10 分钟,epoch 2 的时间为 3 小时。 I want to know what can be the problem我想知道可能是什么问题

Here are some general things that can reduce program speed:以下是一些可以降低程序速度的一般事项:

  • CPU/GPU used by another program另一个程序使用的 CPU/GPU
  • memory swap: your computer moves things from RAM to disk because there is not enough RAM.内存交换:您的计算机将内容从 RAM 移动到磁盘,因为 RAM 不足。 It might be because in your script you try to keep everything in memory (like a list of previous outputs, maybe even with their gradients), or because another program also started to use much RAM.这可能是因为在您的脚本中,您试图将所有内容都保存在内存中(例如之前输出的列表,甚至可能带有它们的渐变),或者因为另一个程序也开始使用大量 RAM。
  • computer heat (maybe it got hot after the first epoch)计算机发热(可能在第一个纪元之后变热了)
  • battery saving (possible if it's a laptop and you unplugged it)省电(如果是笔记本电脑并且您拔掉了电源,则可能)

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM