Movinet fails in GPU mode
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 25/100
- issue の種類
- バグ
- 明瞭さ
- 説明が足りない
- 活発さ
- 停滞
- 技術スタック
- javascript, nodejs, tensorflow
調査の方向性
報告された失敗が発生している src/movinet/MovinetModel.js から始め、GPU を正常に検出する bin/node src/test_gputensorflow.js と比較してください。TensorFlow.js の GPU バックエンドのスタックトレースを確認し、MoViNet の推論で登録済みのプラットフォームを見つけられない理由を特定してください。GPU モードの動画分類が報告されたエラーなしで進行すれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Which version of recognize are you using?
6.11
Enabled Modes
Object recognition, Face recognition, Video recognition, Music recognition
TensorFlow mode
GPU mode
Downstream App
Memories App
Which Nextcloud version do you have installed?
28.0.4.1
Which Operating system do you have installed?
Ubuntu 20.0.4.4
Which database are you running Nextcloud on?
Postgres
Which Docker container are you using to run Nextcloud? (if applicable)
28.0.4.1
How much RAM does your server have?
32
What processor Architecture does your CPU have?
x86_64
Describe the Bug
Seems like after upgrade to NC 28.0.4.1 & Recognize 6.1.1 it started to report below.
It seems like it launches process, based on looking at nvidia-smi and it stays there though doesn't seem like it puts a load on GPU.
Classifier process output: Error: Session fail to run with error: 2 root error(s) found.
(0) NOT_FOUND: could not find registered platform with id: 0x7fd379c7fae4
\t [[{{node movinet_classifier/movinet/stem/stem/conv3d/StatefulPartitionedCall}}]]
\t [[StatefulPartitionedCall/_1555]]
(1) NOT_FOUND: could not find registered platform with id: 0x7fd379c7fae4
\t [[{{node movinet_classifier/movinet/stem/stem/conv3d/StatefulPartitionedCall}}]]
0 successful operations.
0 derived errors ignored.
at NodeJSKernelBackend.runSavedModel (/var/www/html/custom_apps/recognize/node_modules/@tensorflow/tfjs-node-gpu/dist/nodejs_kernel_backend.js:461:43)
at TFSavedModel.predict (/var/www/html/custom_apps/recognize/node_modules/@tensorflow/tfjs-node-gpu/dist/saved_model.js:341:43)
at MovinetModel.predict (/var/www/html/custom_apps/recognize/src/movinet/MovinetModel.js:46:21)
at /var/www/html/custom_apps/recognize/src/movinet/MovinetModel.js:95:24
at /var/www/html/custom_apps/recognize/node_modules/@tensorflow/tfjs-core/dist/tf-core.node.js:4559:22
at Engine.scopedRun (/var/www/html/custom_apps/recognize/node_modules/@tensorflow/tfjs-core/dist/tf-core.node.js:4569:23)
at Engine.tidy (/var/www/html/custom_apps/recognize/node_modules/@tensorflow/tfjs-core/dist/tf-core.node.js:4558:21)
at Object.tidy (/var/www/html/custom_apps/recognize/node_modules/@tensorflow/tfjs-core/dist/tf-core.node.js:8291:19)
at MovinetModel.inference (/var/www/html/custom_apps/recognize/src/movinet/MovinetModel.js:92:21)
at runMicrotasks (<anonymous>)
At the same time it seems to have all what's needed (btw. it didn't raise this error prior upgrades):
./bin/node src/test_gputensorflow.js
2024-04-07 21:37:06.584377: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX2 FMA
To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags.
2024-04-07 21:37:06.593729: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:975] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2024-04-07 21:37:06.636405: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:975] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2024-04-07 21:37:06.636674: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:975] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2024-04-07 21:37:07.109813: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:975] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2024-04-07 21:37:07.110064: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:975] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2024-04-07 21:37:07.110196: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:975] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
2024-04-07 21:37:07.110354: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1532] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 3393 MB memory: -> device: 0, name: Quadro M1200, pci bus id: 0000:01:00.0, compute capability: 5.0
And models seem to be there too:
du -shx ./models/*
50M ./models/efficientnet_lite4
794M ./models/efficientnetv2
22M ./models/landmarks_africa
41M ./models/landmarks_asia
41M ./models/landmarks_europe
41M ./models/landmarks_north_america
31M ./models/landmarks_oceania
41M ./models/landmarks_south_america
47M ./models/movinet-a3
31M ./models/musicnn
Unsure where to look for an issue.
It is after ffmpeg finishes its job.
Thanks!
Expected Behavior
Would just proceed to classify.
To Reproduce
Unsure.
Debug log
No response
- 主要言語
- PHP
- スター
- 699
- フォーク
- 68
- 平均マージ
- 18時間 34分
- マージ済み PR(30日)
- 2
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
nextcloud/recognize のほかの issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
-
insertDeletion looks up recognize_fs_deletions by node_id alone, which no index covers — bulk removals scan the whole table per file対応中かも @marcelklehr が 6 日前に担当しました。 オープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
-
enhancement
難易度 3/5 1〜2日 初心者へのやさしさ 65/100
-
enhancement
難易度 4/5 3〜5日 初心者へのやさしさ 45/100
-
Recognize maintenance routine unnecessarily redownloads Tensorflow CPU and GPU models every time再び着手できるかも @marcelklehr が 187 日前に担当しましたが、オープン中のプルリクエストはありません。 オープンbug
nextcloud/recognize の issue をすべて見る
似ている issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
awslabs/aidlc-workflows#1879 ·
メンテナーはふだん 1 日以内に返信
-
bug customer-reported
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
MagnaCapax/PMSS#1011 ·
メンテナーはふだん 5 日以内に返信
-
Talk Review
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
socallinuxexpo/scale-drupal#351 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
code4romania/cpc#47 ·
メンテナーはふだん 1 日以内に返信
-
Bug Status: Needs Triage
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信