Перейти до основного вмісту

Запуск лише MLA з тензорами INT8

ПолеЗначення
КатегоріяСпівпроцесинг PCIe
СкладністьСередній
Орієнтовний час читання15 хвилин
МіткиPCIe, MLA, INT8, quantization, tensor

Одна програма квантує зображення на хості, виконує маршрут лише MLA, деквантує голови та звіряє їх із типовим маршрутом на тій самій черзі.

Покроковий огляд​

Перегляньте контракт лише MLA​

Створіть Model з увімкненим mla_only. Тепер info() повідомляє входи та виходи INT8, кожен із quant.scale та quant.zero_point. Еталонна модель має один вхід, images, тензор INT8 [640, 640, 3] HWC.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

Квантуйте на хості​

Еталонна модель очікує одне зображення RGB з пікселями в діапазоні [0, 1]. Змініть розмір зображення до 640x640, перетворіть BGR на RGB, поділіть на 255 і застосуйте формулу квантування з параметрами входу. Ця попередня обробка належить моделі, а не контракту архіву: іншій моделі потрібен власний рецепт. Збережіть деквантовані значення тих самих кодів: це точний вхід FP32, який потрібен типовому маршруту для порівняння за однакових умов.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

Виконайте маршрут INT8​

build() запускає конвеєр на карті, що містить лише MLA. run() приймає тензори INT8 і повертає по одному щільному тензору INT8 на кожен вихід у порядку та з іменами, які повідомив info().outputs. Маршрут відхиляє будь-який інший dtype; надсилання FP32 завершується помилкою замість квантування на карті.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

Деквантуйте та порівняйте​

Створіть другу Model без mla_only і надішліть деквантовані значення FP32 через типовий маршрут. Деквантуйте голови INT8 за параметрами кожного виходу та виведіть найбільше відхилення для кожної голови в одиницях її scale. Обидва маршрути виконують ту саму програму MLA з тими самими кодами, тому похибка дорівнює нулю.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

Запуск​

Установіть пакет хоста PCIe та завантажте набір навчальних матеріалів, як описано в розділі Налаштування навчальних матеріалів.

Посібник виконує архів YOLO26n INT8 з Model Zoo, скомпільований для прямого входу та виходу MLA. Завантажте його в корінь розпакованих додаткових матеріалів PCIe:

sima-cli download https://docs.sima.ai/pkg_downloads/SDK2.1.3/models/modalix/yolo26-detection/yolo26n-det-int8-b1.tar.gz

Інші архіви підходять, якщо їх скомпільовано з tessellate_parameters у Model SDK з enable_mla=True, розкладкою DRAM HWC на кожному вході та HWC16 на кожному виході. Архіви, що натомість виконують теселяцію на CVU, як-от збірка yolo_v8s з Model Zoo, відхиляються, коли ввімкнено mla_only:

mla_only does not support stage 'tessellate_quantize_0_MLA_0/...' (tess)

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

C++ (build from source):

./build.sh --target tutorial_027_run_mla_only_int8
./build/tutorials-standalone/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

Обидві версії виводять контракт і нульове відхилення для кожної голови:

MLA-only contract:
input images INT8 [640, 640, 3] scale=0.00390434 zero_point=-128
output bbox_0 INT8 [80, 80, 4] scale=0.0302856 zero_point=-117
...
Dequantized MLA-only outputs vs the default route (error in scale units):
bbox_0 [80, 80, 4] max_err=0.0000
...
[OK] 027_run_mla_only_int8

Типово використовуються карта 0 і черга 0. Передавайте --card N лише за використання іншої карти.

На практиці​

Для архіву INT8 умикайте mla_only, коли застосунок сам відповідає за квантування: він уже отримує INT8 із сенсора або попередньої моделі, йому потрібні сирі голови INT8 для власної постобробки, або він хоче прибрати етапи квантування й деквантування із затримки на боці карти. Читайте кожен scale і zero point з model.info(); ніколи не копіюйте їх з іншої збірки моделі.

Маршрут працює за принципом «усе або нічого». Кожен вхід має відповідати dtype, формі та розміру в байтах свого запису info().inputs, а попередню обробку зображень чи декодування рамок не можна поєднувати з mla_only. Залишайте типовий маршрут, коли хост має дані FP32 і не потребує виконувати перетворення самостійно.

Для діагностики розгортання перейдіть до робочого процесу моделі PCIe.

Повний початковий код​

Показати повні програми
pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
// Quantize on the host, run only the MLA over PCIe, and dequantize the INT8 results.
//
// Usage:
// tutorial_027_run_mla_only_int8 --model yolo26n-det-int8-b1.tar.gz [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <opencv2/imgcodecs.hpp>
#include <opencv2/imgproc.hpp>

#include <algorithm>
#include <cmath>
#include <cstdint>
#include <cstdlib>
#include <filesystem>
#include <iomanip>
#include <iostream>
#include <stdexcept>
#include <string>
#include <utility>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kRunTimeoutMs = 30000;
constexpr double kMaxErrorScales = 0.05; // one twentieth of a quantization step
constexpr char kImagePath[] = "share/sima-pcie-host/tutorials/assets/street-scene.png";

struct Args {
std::string model;
int card_id = 0;
};

std::string require_value(int argc, char** argv, int& index, const char* option) {
if (index + 1 >= argc) {
throw std::runtime_error(std::string("missing value for ") + option);
}
return argv[++index];
}

Args parse_args(int argc, char** argv) {
Args args;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--model") {
args.model = require_value(argc, argv, index, "--model");
} else if (arg == "--card") {
args.card_id = std::stoi(require_value(argc, argv, index, "--card"));
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " --model yolo26n-det-int8-b1.tar.gz [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown argument: " + arg);
}
}
if (args.model.empty()) {
throw std::runtime_error("--model is required: yolo26n-det-int8-b1.tar.gz from the Model Zoo");
}
return args;
}

std::string shape_string(const std::vector<std::int64_t>& shape) {
std::string text = "[";
for (std::size_t index = 0; index < shape.size(); ++index) {
text += (index == 0 ? "" : ", ") + std::to_string(shape[index]);
}
return text + "]";
}

std::size_t element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const auto dim : shape) {
count *= static_cast<std::size_t>(dim);
}
return count;
}

// The MLA-only contract quantizes every tensor with one scale and one zero point:
// x = (q - zero_point) * scale
// q = clamp(round(x / scale) + zero_point, -128, 127)
const pcie::QuantParams& require_quant(const pcie::TensorInfo& info) {
if (!info.quant.has_value()) {
throw std::runtime_error("tensor '" + info.name + "' publishes no quantization parameters");
}
return *info.quant;
}

std::int8_t quantize(const float value, const pcie::QuantParams& quant) {
const float code = std::nearbyint(value / quant.scale) + static_cast<float>(quant.zero_point);
return static_cast<std::int8_t>(std::clamp(code, -128.0F, 127.0F));
}

float dequantize(const std::int8_t code, const pcie::QuantParams& quant) {
return static_cast<float>(static_cast<std::int32_t>(code) - quant.zero_point) * quant.scale;
}

const std::int8_t* int8_data(const pcie::Tensor& tensor) {
return reinterpret_cast<const std::int8_t*>(static_cast<const std::uint8_t*>(tensor.data) +
tensor.byte_offset);
}

// Preprocessing of the reference model: one RGB HWC image with pixels in [0, 1]. Resize, BGR to
// RGB, divide by 255, then quantize with the ingress parameters. Another model needs its own
// recipe.
std::vector<std::int8_t> quantize_image(const cv::Mat& bgr, const pcie::TensorInfo& input) {
if (input.shape.size() != 3 || input.shape[2] != 3) {
throw std::runtime_error("the reference model takes a three-channel HWC input, got " +
shape_string(input.shape));
}
const pcie::QuantParams& quant = require_quant(input);
cv::Mat rgb;
cv::resize(bgr, rgb,
cv::Size(static_cast<int>(input.shape[1]), static_cast<int>(input.shape[0])));
cv::cvtColor(rgb, rgb, cv::COLOR_BGR2RGB);
std::vector<std::int8_t> codes(rgb.total() * rgb.channels());
for (std::size_t index = 0; index < codes.size(); ++index) {
codes[index] = quantize(static_cast<float>(rgb.data[index] / 255.0), quant);
}
return codes;
}

} // namespace

int main(int argc, char** argv) {
try {
const Args args = parse_args(argc, argv);
if (!std::filesystem::is_regular_file(args.model)) {
throw std::runtime_error("model does not exist: " + args.model);
}
const cv::Mat image = cv::imread(kImagePath, cv::IMREAD_COLOR);
if (image.empty()) {
throw std::runtime_error(std::string("OpenCV could not decode: ") + kImagePath);
}
pcie::ConnectionOptions connection;
connection.card_id = args.card_id;

// CORE LOGIC
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

std::cout << "[OK] 027_run_mla_only_int8\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

Джерело​