Tensorflow-serving docker 容器添加了 GPU 設備，但 GPU 的利用率為 0%

Question

嗨，我在使用 dockerized TF Serving 時遇到問題，但沒有使用我的 GPU。

它將 GPU 添加為設備 0，為其分配內存，然后將 ML 模型加載到 CPU 設備內存中，並僅使用 CPU 運行推理。 nvidia-smi 上的 GPU-util 永遠不會離開 0%。

誰能幫我弄清楚為什么會這樣，應該改變什么？

設置：

操作系統： Amazon/Deep Learning AMI (Ubuntu 18.04) on EC2 g4dn.xlarge

GPU：特斯拉T4

模型：來自Huggingface 的預訓練gpt2-xl 張量流，我將其凍結為 SavedModel 並上傳到 S3。

Docker：隨附深度學習 AMI。 我已經檢查並確認 nvidia-smi 運行容器化，所以這不是 nvidia+docker 問題。

TF Serving：我使用以下 Dockerfile 來拉取最新的 gpu 鏡像，並在構建時將模型直接下載到其中：

FROM tensorflow/serving:latest-gpu

RUN apt-get update

ENV TZ=America
RUN ln -snf /usr/share/zoneinfo/$TZ /etc/localtime && echo $TZ > /etc/timezone
RUN apt-get update
RUN apt-get install -y awscli

ENV AWS_ACCESS_KEY_ID=...
ENV AWS_SECRET_ACCESS_KEY=...

ARG model_name
ENV MODEL_NAME=$model_name

# Use AWS CLI to download the SavedModel into the docker container from S3 bucket
RUN aws s3 cp s3://v3-models/models/pretrained_tf_serving/${MODEL_NAME} /models/${MODEL_NAME} --recursive

EXPOSE 8500

我使用以下命令構建並運行上述 Dockerfile：

#!/bin/bash

# first build the image with the model_name arg, and tag it as xl-serving
docker build -t xl-serving --build-arg model_name=gpt2-xl ../../model_server

# then run it with gpus, exposing gRPC port
docker run -it --rm --gpus all --runtime nvidia -p 8500:8500 xl-serving

運行服務容器會打印此輸出。 請注意，添加了 GPU。

2020-11-06 05:25:34.671071: I tensorflow_serving/model_servers/server.cc:87] Building single TensorFlow model file config:  model_name: gpt2-xl model_base_path: /models/gpt2-xl
2020-11-06 05:25:34.671274: I tensorflow_serving/model_servers/server_core.cc:464] Adding/updating models.
2020-11-06 05:25:34.671295: I tensorflow_serving/model_servers/server_core.cc:575]  (Re-)adding model: gpt2-xl
2020-11-06 05:25:34.771644: I tensorflow_serving/core/basic_manager.cc:739] Successfully reserved resources to load servable {name: gpt2-xl version: 1}
2020-11-06 05:25:34.771673: I tensorflow_serving/core/loader_harness.cc:66] Approving load for servable version {name: gpt2-xl version: 1}
2020-11-06 05:25:34.771687: I tensorflow_serving/core/loader_harness.cc:74] Loading servable version {name: gpt2-xl version: 1}
2020-11-06 05:25:34.771724: I external/org_tensorflow/tensorflow/cc/saved_model/reader.cc:31] Reading SavedModel from: /models/gpt2-xl/1
2020-11-06 05:25:35.222512: I external/org_tensorflow/tensorflow/cc/saved_model/reader.cc:54] Reading meta graph with tags { serve }
2020-11-06 05:25:35.222545: I external/org_tensorflow/tensorflow/cc/saved_model/loader.cc:234] Reading SavedModel debug info (if present) from: /models/gpt2-xl/1
2020-11-06 05:25:35.222672: I external/org_tensorflow/tensorflow/core/platform/cpu_feature_guard.cc:142] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN)to use the following CPU instructions in performance-critical operations:  AVX2 AVX512F FMA
To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags.
2020-11-06 05:25:35.223994: I external/org_tensorflow/tensorflow/stream_executor/platform/default/dso_loader.cc:48] Successfully opened dynamic library libcuda.so.1
2020-11-06 05:25:35.262238: I external/org_tensorflow/tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:982] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2020-11-06 05:25:35.263132: I external/org_tensorflow/tensorflow/core/common_runtime/gpu/gpu_device.cc:1716] Found device 0 with properties: 
pciBusID: 0000:00:1e.0 name: Tesla T4 computeCapability: 7.5
coreClock: 1.59GHz coreCount: 40 deviceMemorySize: 14.75GiB deviceMemoryBandwidth: 298.08GiB/s
2020-11-06 05:25:35.263149: I external/org_tensorflow/tensorflow/stream_executor/platform/default/dlopen_checker_stub.cc:25] GPU libraries are statically linked, skip dlopen check.
2020-11-06 05:25:35.263236: I external/org_tensorflow/tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:982] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2020-11-06 05:25:35.264122: I external/org_tensorflow/tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:982] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2020-11-06 05:25:35.264948: I external/org_tensorflow/tensorflow/core/common_runtime/gpu/gpu_device.cc:1858] Adding visible gpu devices: 0
2020-11-06 05:25:36.185140: I external/org_tensorflow/tensorflow/core/common_runtime/gpu/gpu_device.cc:1257] Device interconnect StreamExecutor with strength 1 edge matrix:
2020-11-06 05:25:36.185165: I external/org_tensorflow/tensorflow/core/common_runtime/gpu/gpu_device.cc:1263]      0 
2020-11-06 05:25:36.185171: I external/org_tensorflow/tensorflow/core/common_runtime/gpu/gpu_device.cc:1276] 0:   N 
2020-11-06 05:25:36.185334: I external/org_tensorflow/tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:982] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2020-11-06 05:25:36.186222: I external/org_tensorflow/tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:982] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2020-11-06 05:25:36.187046: I external/org_tensorflow/tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:982] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2020-11-06 05:25:36.187852: I external/org_tensorflow/tensorflow/core/common_runtime/gpu/gpu_device.cc:1402] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:0 with 13896 MB memory) -> physical GPU (device: 0, name: Tesla T4, pci bus id: 0000:00:1e.0, compute capability: 7.5)
2020-11-06 05:25:37.279837: I external/org_tensorflow/tensorflow/cc/saved_model/loader.cc:199] Restoring SavedModel bundle.
2020-11-06 05:25:56.154008: I external/org_tensorflow/tensorflow/cc/saved_model/loader.cc:183] Running initialization op on SavedModel bundle at path: /models/gpt2-xl/1
2020-11-06 05:25:57.551535: I external/org_tensorflow/tensorflow/cc/saved_model/loader.cc:303] SavedModel load for tags { serve }; Status: success: OK. Took 22777844 microseconds.
2020-11-06 05:25:57.832736: I tensorflow_serving/servables/tensorflow/saved_model_warmup_util.cc:59] No warmup data file found at /models/gpt2-xl/1/assets.extra/tf_serving_warmup_requests
2020-11-06 05:25:57.835030: I tensorflow_serving/core/loader_harness.cc:87] Successfully loaded servable version {name: gpt2-xl version: 1}
2020-11-06 05:25:57.838329: I tensorflow_serving/model_servers/server.cc:367] Running gRPC ModelServer at 0.0.0.0:8500 ...
[warn] getaddrinfo: address family for nodename not supported
2020-11-06 05:25:57.840415: I tensorflow_serving/model_servers/server.cc:387] Exporting HTTP/REST API at:localhost:8501 ...
[evhttp_server.cc : 238] NET_LOG: Entering the event loop ...

然后，我使用單個非批處理 gRPC 調用訪問了該服務器。 它將成功運行並返回正確的 GPT2 輸出。 但是，只要在 CPU 上進行相同的設置就需要花費很長時間。 htop 顯示 8gb 的 ram（gpt2-xl 模型大小）已加載到 CPU 機器中。 然后它會顯示 TF Serving 進程正在運行，並最大限度地使用一兩個 CPU 內核。 它似乎只在 CPU 上運行。

這就是調用運行時nvidia-smi 的樣子。 注意分配的內存和 0% GPU-Util：

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 450.80.02    Driver Version: 450.80.02    CUDA Version: 11.0     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Tesla T4            On   | 00000000:00:1E.0 Off |                    0 |
| N/A   36C    P0    26W /  70W |  14240MiB / 15109MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+
                                                                               
+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
|    0   N/A  N/A     13357      C   tensorflow_model_server         14221MiB |
+-----------------------------------------------------------------------------+

我已經在網上搜索過，但找不到任何建議。 我發現最接近的是這個 github 問題： GPU 利用率與 TF 服務 #1440 ，修復對我不起作用。 他們正在處理低 GPU 利用率，我正在處理 0%。

關於這個問題的任何建議？

非常感謝。 這幾天我一直在用頭撞牆，所以我非常感謝你的幫助:)

更新#1：

我已經編寫了一個 python 腳本（如下）來使用 tensorflow==2.3.0 來加載模型並運行它。 它在 CUDA=11.0 的 conda 環境中運行。 它成功地在 GPU 上運行推理，並且比我在 TF 服務上得到的快 15 倍。

import tensorflow as tf
import numpy as np

model = tf.saved_model.load('/home/ubuntu/models/gpt2-xl/1/')
servable = model.signatures["forward"]

# Create input tensor
tensor_in = tf.constant([[198, 15667,  6530, 25, 29437, 1706, 1610, 977, 948, 33611]])

# Run a loop of 10 inferences on the model, to predict the next 10 tokens.
for i in range(10):
    pred = servable(tensor_in)
    logits = pred['output_0']
    logits = logits[:, -1, :] / 0.8
    next_id = tf.random.categorical(tf.nn.log_softmax(logits, axis=-1), num_samples=1)
    next_id = tf.dtypes.cast(next_id, tf.int32).numpy()
    tensor_in = np.concatenate((tensor_in, next_id), axis=1)

接下來：將嘗試在容器外運行 tf-serving。 更新即將到來...

Answer 1

你是如何保存模型的？ 保存模型時添加clear_devices=True並再次嘗試。

Tensorflow-serving docker 容器添加了 GPU 設備，但 GPU 的利用率為 0%

問題描述

嗨，我在使用 dockerized TF Serving 時遇到問題，但沒有使用我的 GPU。

更新#1：

1 個解決方案

解決方案1
0 2020-11-12 10:14:06

Tensorflow-serving docker 容器添加了 GPU 設備，但 GPU 的利用率為 0%

問題描述

嗨，我在使用 dockerized TF Serving 時遇到問題，但沒有使用我的 GPU。

更新#1：

1 個解決方案

解決方案1 0 2020-11-12 10:14:06

解決方案1
0 2020-11-12 10:14:06