Skip to main content

Wrap External Tensor Memory

Wrap External Tensor Memory — animated walkthrough overview

FieldValue
CategoryPCIe Co-Processing
DifficultyIntermediate
Estimated Read Time15 minutes
LabelsPCIe, C++, tensor, external memory, zero-copy wrapping

The program creates three reusable input slots and submits eight synthetic FP32 frames. A slot returns to the available queue only after the ordered result for that slot is pulled.

Walkthrough​

Inspect the model contract​

Construct the model and read info().inputs before allocating memory. This tutorial uses the single-input YOLOv8s archive and verifies that its reported dtype is FP32 and that its shape accounts for exactly size_bytes bytes.

An external view must match the corresponding TensorInfo dtype, shape, byte size, and name. Do not infer these values from another model build.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

Wrap application-owned memory​

Each ring slot owns a std::shared_ptr<std::vector<float>>. The call to Tensor::from_external() receives the base pointer, complete backing element count, shared owner, model shape, and route name. Because the view is contiguous, the PCIe host can wrap it directly instead of creating a staging allocation.

Keeping only a raw pointer is not sufficient. The shared owner is mandatory because the transport can retain the tensor after push() returns.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

Build the model​

Build after the model contract and ring allocations have been validated. The example sets max_inflight to the ring size so the application and transport have the same explicit bound.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
model.build(kBuildTimeoutMs);

Submit and safely reuse the ring​

Fill an available slot, call push(), and move that slot to the in-flight queue. Do not modify or reuse its storage merely because push() returned. The example calls pull() when no slot is available and returns the oldest slot to the available queue only after its matching ordered result arrives.

Every accepted push is balanced by one pull, including the final drain. A timeout closes the model rather than reusing memory whose request may still be active.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

Run​

Install the PCIe host package and download the tutorial bundle as described in Tutorial Setup. Download YOLOv8s into the extracted PCIe extras root:

sima-cli modelzoo get yolo_v8s
cp /absolute/path/to/downloaded-yolov8s-archive.tar.gz yolo_v8s_mpk.tar.gz
test -f yolo_v8s_mpk.tar.gz

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_028_wrap_external_tensor_memory

C++ (build from source):

./build.sh --target tutorial_028_wrap_external_tensor_memory
./build/tutorials-standalone/tutorial_028_wrap_external_tensor_memory

The default is card 0 and queue 0. Pass --card N only when using another card. A successful run prints:

input=images
ring_slots=3
completed=8
[OK] 028_wrap_external_tensor_memory

In Practice​

The direct host wrapping path requires contiguous storage. A tensor with non-contiguous strides is still accepted when its descriptor is valid, but the host compacts it into a staging allocation. Multiple separately allocated inputs are also packed into staging memory.

For a multi-input model, avoid that staging allocation only when all tensors are consecutive views into one shared packed allocation. Submit them in the order reported by info().inputs, and use each input's name and shape. For example, when both inputs are FP32:

const auto& first = info.inputs.at(0);
const auto& second = info.inputs.at(1);
const std::size_t first_count = first.size_bytes / sizeof(float);
const std::size_t second_count = second.size_bytes / sizeof(float);

auto packed =
std::make_shared<std::vector<float>>(first_count + second_count);
pcie::Tensor input0 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, first.shape, first.name);
pcie::Tensor input1 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, second.shape, second.name,
static_cast<std::int64_t>(first.size_bytes));

model.push({input0, input1});

This optimization removes only the host-side packing copy. PCIe still copies the packed payload into card-owned transport memory before inference.

Full source​

Show the complete source programs
pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
// Submit application-owned tensor memory without a host staging copy.
//
// Usage:
// tutorial_028_wrap_external_tensor_memory [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <algorithm>
#include <cstdint>
#include <cstdlib>
#include <deque>
#include <filesystem>
#include <iostream>
#include <limits>
#include <memory>
#include <stdexcept>
#include <string>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kPullTimeoutMs = 30000;
constexpr std::size_t kRingSlots = 3;
constexpr std::size_t kFrameCount = 8;
constexpr char kModelPath[] = "yolo_v8s_mpk.tar.gz";

int parse_card(const int argc, char** argv) {
int card_id = 0;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--card" && index + 1 < argc) {
card_id = std::stoi(argv[++index]);
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown or incomplete argument: " + arg);
}
}
return card_id;
}

std::size_t checked_element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const std::int64_t dimension : shape) {
if (dimension <= 0 ||
count > std::numeric_limits<std::size_t>::max() / static_cast<std::size_t>(dimension)) {
throw std::runtime_error("model input has an invalid or overflowing shape");
}
count *= static_cast<std::size_t>(dimension);
}
return count;
}

void validate_fp32_input(const pcie::TensorInfo& input) {
if (input.dtype != "FP32" && input.dtype != "FLOAT32") {
throw std::runtime_error("tutorial requires an FP32 model input, got " + input.dtype);
}
const std::size_t element_count = checked_element_count(input.shape);
if (element_count > std::numeric_limits<std::size_t>::max() / sizeof(float) ||
element_count * sizeof(float) != input.size_bytes) {
throw std::runtime_error("model input shape and byte size are inconsistent");
}
}

struct InputSlot {
std::shared_ptr<std::vector<float>> storage;
pcie::Tensor tensor;
};

InputSlot make_input_slot(const pcie::TensorInfo& input) {
auto storage = std::make_shared<std::vector<float>>(checked_element_count(input.shape), 0.0F);
pcie::Tensor tensor = pcie::Tensor::from_external(storage->data(), storage->size(), storage,
input.shape, input.name);
return {.storage = std::move(storage), .tensor = std::move(tensor)};
}

std::vector<InputSlot> make_input_ring(const pcie::TensorInfo& input) {
std::vector<InputSlot> slots;
slots.reserve(kRingSlots);
for (std::size_t index = 0; index < kRingSlots; ++index) {
slots.push_back(make_input_slot(input));
}
return slots;
}

std::size_t run_input_ring(pcie::Model& model, std::vector<InputSlot>& slots) {
std::deque<std::size_t> available;
std::deque<std::size_t> in_flight;
for (std::size_t index = 0; index < slots.size(); ++index) {
available.push_back(index);
}

std::size_t completed = 0;
const auto complete_oldest = [&] {
auto outputs = model.pull(kPullTimeoutMs);
if (!outputs) {
throw std::runtime_error("timed out waiting for an external-memory submission");
}
if (outputs->empty() || in_flight.empty()) {
throw std::runtime_error("received an invalid external-memory completion");
}
available.push_back(in_flight.front());
in_flight.pop_front();
++completed;
};

for (std::size_t frame = 0; frame < kFrameCount; ++frame) {
if (available.empty()) {
complete_oldest();
}

const std::size_t slot_index = available.front();
available.pop_front();
InputSlot& slot = slots[slot_index];

// A slot is writable only while it is not in flight.
std::fill(slot.storage->begin(), slot.storage->end(), static_cast<float>(frame % 10U) / 10.0F);
if (!model.push(slot.tensor)) {
throw std::runtime_error("push rejected frame " + std::to_string(frame));
}
in_flight.push_back(slot_index);
}

while (!in_flight.empty()) {
complete_oldest();
}
return completed;
}

} // namespace

int main(int argc, char** argv) {
try {
const int card_id = parse_card(argc, argv);
if (!std::filesystem::is_regular_file(kModelPath)) {
throw std::runtime_error(std::string("model does not exist: ") + kModelPath);
}

pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

model.build(kBuildTimeoutMs);

std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

std::cout << "input=" << info.inputs.front().name << '\n'
<< "ring_slots=" << slots.size() << '\n'
<< "completed=" << completed << '\n'
<< "[OK] 028_wrap_external_tensor_memory\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

Source​