본문으로 건너뛰기

외부 텐서 메모리 래핑

필드값
범주PCIe 코프로세싱
난이도중급
예상 소요 시간15분
레이블PCIe, C++, tensor, external memory, zero-copy wrapping

프로그램은 재사용 가능한 입력 슬롯 3개를 만들고 합성 FP32 프레임 8개를 제출합니다. 슬롯은 해당 슬롯의 순서가 보장된 결과를 가져온 후에만 사용 가능 큐로 돌아갑니다.

둘러보기​

모델 계약 검사​

메모리를 할당하기 전에 모델을 생성하고 info().inputs를 읽습니다. 이 튜토리얼은 단일 입력 YOLOv8s 아카이브를 사용하며 보고된 dtype이 FP32이고 형상이 정확히 size_bytes 바이트를 차지하는지 확인합니다.

외부 뷰는 해당 TensorInfo의 dtype, 형상, 바이트 크기, 이름과 일치해야 합니다. 다른 모델 빌드에서 이 값을 추정하지 마십시오.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

애플리케이션 소유 메모리 래핑​

각 링 슬롯은 std::shared_ptr<std::vector<float>>를 소유합니다. Tensor::from_external()에는 기본 포인터, 전체 기반 요소 수, 공유 소유자, 모델 형상, 라우트 이름을 전달합니다. 뷰가 연속적이므로 PCIe 호스트는 스테이징 할당을 만들지 않고 직접 래핑할 수 있습니다.

원시 포인터만 유지하는 것으로는 충분하지 않습니다. 전송 계층이 push() 반환 후에도 텐서를 보유할 수 있으므로 공유 소유자가 필수입니다.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

모델 빌드​

모델 계약과 링 할당을 검증한 후 빌드합니다. 예제는 max_inflight를 링 크기로 설정하여 애플리케이션과 전송 계층에 동일한 명시적 한도를 적용합니다.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
model.build(kBuildTimeoutMs);

링 제출 및 안전한 재사용​

사용 가능한 슬롯을 채우고 push()를 호출한 뒤 해당 슬롯을 처리 중 큐로 이동합니다. push()가 반환했다는 이유만으로 저장 공간을 수정하거나 재사용하지 마십시오. 사용 가능한 슬롯이 없으면 예제는 pull()을 호출하고 일치하는 순서의 결과가 도착한 후에만 가장 오래된 슬롯을 사용 가능 큐로 되돌립니다.

마지막 드레인을 포함하여 수락된 모든 푸시마다 한 번씩 풀합니다. 시간 초과가 발생하면 아직 활성 상태일 수 있는 요청의 메모리를 재사용하지 않고 모델을 닫습니다.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

실행​

PCIe 호스트 패키지를 설치하고 튜토리얼 설정에 설명된 대로 튜토리얼 번들을 다운로드합니다. 압축을 푼 PCIe extras 루트에 YOLOv8s를 다운로드합니다.

sima-cli modelzoo get yolo_v8s
cp /absolute/path/to/downloaded-yolov8s-archive.tar.gz yolo_v8s_mpk.tar.gz
test -f yolo_v8s_mpk.tar.gz

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_028_wrap_external_tensor_memory

C++ (build from source):

./build.sh --target tutorial_028_wrap_external_tensor_memory
./build/tutorials-standalone/tutorial_028_wrap_external_tensor_memory

기본값은 카드 0과 큐 0입니다. 다른 카드를 사용할 때만 --card N을 전달하십시오. 성공하면 다음이 출력됩니다.

input=images
ring_slots=3
completed=8
[OK] 028_wrap_external_tensor_memory

실전 활용​

직접 호스트 래핑 경로에는 연속 저장 공간이 필요합니다. 불연속 스트라이드가 있는 텐서도 설명자가 유효하면 허용되지만, 호스트가 스테이징 할당으로 압축합니다. 별도로 할당된 여러 입력도 스테이징 메모리로 패킹됩니다.

다중 입력 모델에서는 모든 텐서가 하나의 공유 패킹 할당에 대한 연속 뷰일 때만 이 스테이징 할당을 피할 수 있습니다. info().inputs가 보고한 순서로 제출하고 각 입력의 이름과 형상을 사용하십시오. 예를 들어 두 입력이 모두 FP32인 경우는 다음과 같습니다.

const auto& first = info.inputs.at(0);
const auto& second = info.inputs.at(1);
const std::size_t first_count = first.size_bytes / sizeof(float);
const std::size_t second_count = second.size_bytes / sizeof(float);

auto packed =
std::make_shared<std::vector<float>>(first_count + second_count);
pcie::Tensor input0 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, first.shape, first.name);
pcie::Tensor input1 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, second.shape, second.name,
static_cast<std::int64_t>(first.size_bytes));

model.push({input0, input1});

이 최적화는 호스트 측 패킹 복사만 제거합니다. PCIe는 추론 전에 패킹된 페이로드를 카드 소유 전송 메모리로 계속 복사합니다.

전체 소스​

전체 소스 프로그램 표시
pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
// Submit application-owned tensor memory without a host staging copy.
//
// Usage:
// tutorial_028_wrap_external_tensor_memory [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <algorithm>
#include <cstdint>
#include <cstdlib>
#include <deque>
#include <filesystem>
#include <iostream>
#include <limits>
#include <memory>
#include <stdexcept>
#include <string>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kPullTimeoutMs = 30000;
constexpr std::size_t kRingSlots = 3;
constexpr std::size_t kFrameCount = 8;
constexpr char kModelPath[] = "yolo_v8s_mpk.tar.gz";

int parse_card(const int argc, char** argv) {
int card_id = 0;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--card" && index + 1 < argc) {
card_id = std::stoi(argv[++index]);
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown or incomplete argument: " + arg);
}
}
return card_id;
}

std::size_t checked_element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const std::int64_t dimension : shape) {
if (dimension <= 0 ||
count > std::numeric_limits<std::size_t>::max() / static_cast<std::size_t>(dimension)) {
throw std::runtime_error("model input has an invalid or overflowing shape");
}
count *= static_cast<std::size_t>(dimension);
}
return count;
}

void validate_fp32_input(const pcie::TensorInfo& input) {
if (input.dtype != "FP32" && input.dtype != "FLOAT32") {
throw std::runtime_error("tutorial requires an FP32 model input, got " + input.dtype);
}
const std::size_t element_count = checked_element_count(input.shape);
if (element_count > std::numeric_limits<std::size_t>::max() / sizeof(float) ||
element_count * sizeof(float) != input.size_bytes) {
throw std::runtime_error("model input shape and byte size are inconsistent");
}
}

struct InputSlot {
std::shared_ptr<std::vector<float>> storage;
pcie::Tensor tensor;
};

InputSlot make_input_slot(const pcie::TensorInfo& input) {
auto storage = std::make_shared<std::vector<float>>(checked_element_count(input.shape), 0.0F);
pcie::Tensor tensor = pcie::Tensor::from_external(storage->data(), storage->size(), storage,
input.shape, input.name);
return {.storage = std::move(storage), .tensor = std::move(tensor)};
}

std::vector<InputSlot> make_input_ring(const pcie::TensorInfo& input) {
std::vector<InputSlot> slots;
slots.reserve(kRingSlots);
for (std::size_t index = 0; index < kRingSlots; ++index) {
slots.push_back(make_input_slot(input));
}
return slots;
}

std::size_t run_input_ring(pcie::Model& model, std::vector<InputSlot>& slots) {
std::deque<std::size_t> available;
std::deque<std::size_t> in_flight;
for (std::size_t index = 0; index < slots.size(); ++index) {
available.push_back(index);
}

std::size_t completed = 0;
const auto complete_oldest = [&] {
auto outputs = model.pull(kPullTimeoutMs);
if (!outputs) {
throw std::runtime_error("timed out waiting for an external-memory submission");
}
if (outputs->empty() || in_flight.empty()) {
throw std::runtime_error("received an invalid external-memory completion");
}
available.push_back(in_flight.front());
in_flight.pop_front();
++completed;
};

for (std::size_t frame = 0; frame < kFrameCount; ++frame) {
if (available.empty()) {
complete_oldest();
}

const std::size_t slot_index = available.front();
available.pop_front();
InputSlot& slot = slots[slot_index];

// A slot is writable only while it is not in flight.
std::fill(slot.storage->begin(), slot.storage->end(), static_cast<float>(frame % 10U) / 10.0F);
if (!model.push(slot.tensor)) {
throw std::runtime_error("push rejected frame " + std::to_string(frame));
}
in_flight.push_back(slot_index);
}

while (!in_flight.empty()) {
complete_oldest();
}
return completed;
}

} // namespace

int main(int argc, char** argv) {
try {
const int card_id = parse_card(argc, argv);
if (!std::filesystem::is_regular_file(kModelPath)) {
throw std::runtime_error(std::string("model does not exist: ") + kModelPath);
}

pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

model.build(kBuildTimeoutMs);

std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

std::cout << "input=" << info.inputs.front().name << '\n'
<< "ring_slots=" << slots.size() << '\n'
<< "completed=" << completed << '\n'
<< "[OK] 028_wrap_external_tensor_memory\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

소스​