メインコンテンツまでスキップ

外部テンソルメモリをラップする

項目値
カテゴリPCIe コプロセッシング
難易度中級
推定所要時間15分
ラベルPCIe, C++, tensor, external memory, zero-copy wrapping

プログラムは再利用可能な3つの入力スロットを作成し、8つの合成FP32フレームを送信します。スロットは、そのスロットに対応する順序付き結果がプルされた後でのみ、利用可能なキューに戻ります。

ウォークスルー​

モデルのコントラクトを確認する​

メモリを割り当てる前にモデルを構築し、info().inputsを読み取ります。このチュートリアルは単一入力のYOLOv8sアーカイブを使用し、報告されたdtypeがFP32であり、その形状が正確にsize_bytesバイトになることを確認します。

外部ビューは、対応するTensorInfoのdtype、形状、バイトサイズ、名前と一致する必要があります。別のモデルビルドからこれらの値を推測しないでください。

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

アプリケーション所有のメモリをラップする​

各リングスロットはstd::shared_ptr<std::vector<float>>を所有します。Tensor::from_external()には、ベースポインタ、完全なバッキング要素数、共有所有者、モデル形状、ルート名を渡します。ビューが連続しているため、PCIeホストはステージング割り当てを作成せずに直接ラップできます。

生ポインタだけを保持するのでは不十分です。トランスポートはpush()が戻った後もテンソルを保持できるため、共有所有者が必須です。

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

モデルをビルドする​

モデルコントラクトとリング割り当てを検証した後でビルドします。この例ではmax_inflightをリングサイズに設定し、アプリケーションとトランスポートに同じ明示的な上限を与えます。

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
model.build(kBuildTimeoutMs);

リングを送信して安全に再利用する​

利用可能なスロットを埋め、push()を呼び出し、そのスロットを処理中キューに移します。push()が戻っただけでストレージを変更または再利用しないでください。利用可能なスロットがない場合、この例はpull()を呼び出し、一致する順序付き結果が届いた後でのみ最も古いスロットを利用可能なキューへ戻します。

最後のドレインを含め、受理されたすべてのプッシュに1回のプルを対応させます。タイムアウト時は、まだアクティブかもしれないリクエストのメモリを再利用せず、モデルを閉じます。

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

実行​

PCIeホストパッケージをインストールし、チュートリアルの設定で説明されているようにチュートリアルバンドルをダウンロードします。抽出したPCIeエクストラのルートへYOLOv8sをダウンロードします。

sima-cli modelzoo get yolo_v8s
cp /absolute/path/to/downloaded-yolov8s-archive.tar.gz yolo_v8s_mpk.tar.gz
test -f yolo_v8s_mpk.tar.gz

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_028_wrap_external_tensor_memory

C++ (build from source):

./build.sh --target tutorial_028_wrap_external_tensor_memory
./build/tutorials-standalone/tutorial_028_wrap_external_tensor_memory

デフォルトはカード0とキュー0です。別のカードを使用する場合にのみ--card Nを渡します。成功すると次のように表示されます。

input=images
ring_slots=3
completed=8
[OK] 028_wrap_external_tensor_memory

実践​

ホストで直接ラップする経路には連続したストレージが必要です。不連続ストライドのテンソルも記述子が有効なら受け付けられますが、ホストはステージング割り当てへコンパクト化します。個別に割り当てられた複数入力もステージングメモリへパックされます。

複数入力モデルでこのステージング割り当てを避けられるのは、すべてのテンソルが1つの共有パック割り当て内の連続したビューである場合だけです。info().inputsが報告する順序で送信し、各入力の名前と形状を使用します。たとえば、両方の入力がFP32の場合は次のようになります。

const auto& first = info.inputs.at(0);
const auto& second = info.inputs.at(1);
const std::size_t first_count = first.size_bytes / sizeof(float);
const std::size_t second_count = second.size_bytes / sizeof(float);

auto packed =
std::make_shared<std::vector<float>>(first_count + second_count);
pcie::Tensor input0 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, first.shape, first.name);
pcie::Tensor input1 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, second.shape, second.name,
static_cast<std::int64_t>(first.size_bytes));

model.push({input0, input1});

この最適化が取り除くのはホスト側のパッキングコピーだけです。PCIeは推論前に、パックされたペイロードをカード所有のトランスポートメモリへ引き続きコピーします。

完全なソース​

完全なソースプログラムを表示
pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
// Submit application-owned tensor memory without a host staging copy.
//
// Usage:
// tutorial_028_wrap_external_tensor_memory [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <algorithm>
#include <cstdint>
#include <cstdlib>
#include <deque>
#include <filesystem>
#include <iostream>
#include <limits>
#include <memory>
#include <stdexcept>
#include <string>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kPullTimeoutMs = 30000;
constexpr std::size_t kRingSlots = 3;
constexpr std::size_t kFrameCount = 8;
constexpr char kModelPath[] = "yolo_v8s_mpk.tar.gz";

int parse_card(const int argc, char** argv) {
int card_id = 0;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--card" && index + 1 < argc) {
card_id = std::stoi(argv[++index]);
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown or incomplete argument: " + arg);
}
}
return card_id;
}

std::size_t checked_element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const std::int64_t dimension : shape) {
if (dimension <= 0 ||
count > std::numeric_limits<std::size_t>::max() / static_cast<std::size_t>(dimension)) {
throw std::runtime_error("model input has an invalid or overflowing shape");
}
count *= static_cast<std::size_t>(dimension);
}
return count;
}

void validate_fp32_input(const pcie::TensorInfo& input) {
if (input.dtype != "FP32" && input.dtype != "FLOAT32") {
throw std::runtime_error("tutorial requires an FP32 model input, got " + input.dtype);
}
const std::size_t element_count = checked_element_count(input.shape);
if (element_count > std::numeric_limits<std::size_t>::max() / sizeof(float) ||
element_count * sizeof(float) != input.size_bytes) {
throw std::runtime_error("model input shape and byte size are inconsistent");
}
}

struct InputSlot {
std::shared_ptr<std::vector<float>> storage;
pcie::Tensor tensor;
};

InputSlot make_input_slot(const pcie::TensorInfo& input) {
auto storage = std::make_shared<std::vector<float>>(checked_element_count(input.shape), 0.0F);
pcie::Tensor tensor = pcie::Tensor::from_external(storage->data(), storage->size(), storage,
input.shape, input.name);
return {.storage = std::move(storage), .tensor = std::move(tensor)};
}

std::vector<InputSlot> make_input_ring(const pcie::TensorInfo& input) {
std::vector<InputSlot> slots;
slots.reserve(kRingSlots);
for (std::size_t index = 0; index < kRingSlots; ++index) {
slots.push_back(make_input_slot(input));
}
return slots;
}

std::size_t run_input_ring(pcie::Model& model, std::vector<InputSlot>& slots) {
std::deque<std::size_t> available;
std::deque<std::size_t> in_flight;
for (std::size_t index = 0; index < slots.size(); ++index) {
available.push_back(index);
}

std::size_t completed = 0;
const auto complete_oldest = [&] {
auto outputs = model.pull(kPullTimeoutMs);
if (!outputs) {
throw std::runtime_error("timed out waiting for an external-memory submission");
}
if (outputs->empty() || in_flight.empty()) {
throw std::runtime_error("received an invalid external-memory completion");
}
available.push_back(in_flight.front());
in_flight.pop_front();
++completed;
};

for (std::size_t frame = 0; frame < kFrameCount; ++frame) {
if (available.empty()) {
complete_oldest();
}

const std::size_t slot_index = available.front();
available.pop_front();
InputSlot& slot = slots[slot_index];

// A slot is writable only while it is not in flight.
std::fill(slot.storage->begin(), slot.storage->end(), static_cast<float>(frame % 10U) / 10.0F);
if (!model.push(slot.tensor)) {
throw std::runtime_error("push rejected frame " + std::to_string(frame));
}
in_flight.push_back(slot_index);
}

while (!in_flight.empty()) {
complete_oldest();
}
return completed;
}

} // namespace

int main(int argc, char** argv) {
try {
const int card_id = parse_card(argc, argv);
if (!std::filesystem::is_regular_file(kModelPath)) {
throw std::runtime_error(std::string("model does not exist: ") + kModelPath);
}

pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

model.build(kBuildTimeoutMs);

std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

std::cout << "input=" << info.inputs.front().name << '\n'
<< "ring_slots=" << slots.size() << '\n'
<< "completed=" << completed << '\n'
<< "[OK] 028_wrap_external_tensor_memory\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

ソース​