본문으로 건너뛰기

INT8 텐서로 MLA만 실행하기

필드값
범주PCIe 코프로세싱
난이도중급
예상 소요 시간15분
레이블PCIe, MLA, INT8, quantization, tensor

하나의 프로그램이 호스트에서 이미지를 양자화하고, MLA 전용 경로를 실행하고, 헤드를 역양자화한 뒤, 같은 큐의 기본 경로와 비교 검증합니다.

둘러보기​

MLA 전용 계약 확인하기​

mla_only를 활성화하여 Model을 생성합니다. 이제 info()는 INT8 입력과 출력을 보고하며, 각각 quant.scale과 quant.zero_point을 가집니다. 참조 모델의 입력은 images 하나이며, INT8 [640, 640, 3] HWC 텐서입니다.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

호스트에서 양자화하기​

참조 모델은 픽셀이 [0, 1] 범위인 RGB 이미지 하나를 기대합니다. 이미지를 640x640으로 크기 조정하고, BGR을 RGB로 변환하고, 255로 나눈 뒤, 진입 매개변수로 양자화 수식을 적용합니다. 이 전처리는 아카이브 계약이 아니라 모델에 속합니다. 다른 모델에는 자체 레시피가 필요합니다. 같은 코드를 역양자화한 값을 보관해 두세요. 이 값이 기본 경로가 동일 조건 비교에 필요로 하는 정확한 FP32 입력입니다.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

INT8 경로 실행하기​

build()는 MLA만 포함하는 카드 파이프라인을 시작합니다. run()은 INT8 텐서를 받아 info().outputs가 보고한 순서와 이름으로 출력마다 하나의 밀집 INT8 텐서를 반환합니다. 이 경로는 다른 dtype을 거부합니다. FP32 푸시는 카드에서 양자화되는 대신 실패합니다.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

역양자화하고 비교하기​

mla_only 없이 두 번째 Model을 생성하고 역양자화된 FP32 값을 기본 경로로 보냅니다. 각 출력의 매개변수로 INT8 헤드를 역양자화하고 헤드별 최대 편차를 해당 헤드의 스케일 단위로 출력합니다. 두 경로 모두 같은 코드로 같은 MLA 프로그램을 실행하므로 오차는 0입니다.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

실행​

튜토리얼 설정에 설명된 대로 PCIe 호스트 패키지를 설치하고 튜토리얼 번들을 다운로드합니다.

이 튜토리얼은 MLA 입출력을 직접 처리하도록 컴파일된 Model Zoo의 YOLO26n INT8 아카이브를 실행합니다. 추출된 PCIe extras 루트에 다운로드하세요.

sima-cli download https://docs.sima.ai/pkg_downloads/SDK2.1.3/models/modalix/yolo26-detection/yolo26n-det-int8-b1.tar.gz

다른 아카이브는 Model SDK의 tessellate_parameters에서 enable_mla=True를 지정하고, 모든 입력에 HWC DRAM 레이아웃을, 모든 출력에 HWC16을 지정하여 컴파일된 경우 조건을 충족합니다. Model Zoo의 yolo_v8s 빌드처럼 대신 CVU에서 테셀레이션을 수행하는 아카이브는 mla_only가 활성화되면 거부됩니다.

mla_only does not support stage 'tessellate_quantize_0_MLA_0/...' (tess)

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

C++ (build from source):

./build.sh --target tutorial_027_run_mla_only_int8
./build/tutorials-standalone/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

두 버전 모두 계약과 모든 헤드에 대한 0의 편차를 출력합니다.

MLA-only contract:
input images INT8 [640, 640, 3] scale=0.00390434 zero_point=-128
output bbox_0 INT8 [80, 80, 4] scale=0.0302856 zero_point=-117
...
Dequantized MLA-only outputs vs the default route (error in scale units):
bbox_0 [80, 80, 4] max_err=0.0000
...
[OK] 027_run_mla_only_int8

기본값은 카드 0과 큐 0입니다. 다른 카드를 사용할 때만 --card N을 전달하세요.

실전 활용​

INT8 아카이브에서는 애플리케이션이 양자화를 소유할 때 mla_only를 활성화하세요. 센서나 앞단 모델에서 이미 INT8을 생성하는 경우, 자체 후처리를 위해 원시 INT8 헤드가 필요한 경우, 또는 카드 측 지연 시간에서 양자화와 역양자화 단계를 제거하려는 경우입니다. 모든 스케일과 영점은 model.info()에서 읽고, 모델의 다른 빌드에서 복사하지 마세요.

이 경로는 전부 아니면 전무입니다. 모든 입력은 해당 info().inputs 항목의 dtype, 형상, 바이트 크기와 일치해야 하며, 이미지 전처리나 박스 디코드를 mla_only와 결합할 수 없습니다. 호스트가 FP32 데이터를 보유하고 변환을 직접 처리할 필요가 없다면 기본 경로를 유지하세요.

배포 진단은 PCIe 모델 워크플로에서 계속하세요.

전체 소스​

전체 소스 프로그램 표시
pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
// Quantize on the host, run only the MLA over PCIe, and dequantize the INT8 results.
//
// Usage:
// tutorial_027_run_mla_only_int8 --model yolo26n-det-int8-b1.tar.gz [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <opencv2/imgcodecs.hpp>
#include <opencv2/imgproc.hpp>

#include <algorithm>
#include <cmath>
#include <cstdint>
#include <cstdlib>
#include <filesystem>
#include <iomanip>
#include <iostream>
#include <stdexcept>
#include <string>
#include <utility>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kRunTimeoutMs = 30000;
constexpr double kMaxErrorScales = 0.05; // one twentieth of a quantization step
constexpr char kImagePath[] = "share/sima-pcie-host/tutorials/assets/street-scene.png";

struct Args {
std::string model;
int card_id = 0;
};

std::string require_value(int argc, char** argv, int& index, const char* option) {
if (index + 1 >= argc) {
throw std::runtime_error(std::string("missing value for ") + option);
}
return argv[++index];
}

Args parse_args(int argc, char** argv) {
Args args;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--model") {
args.model = require_value(argc, argv, index, "--model");
} else if (arg == "--card") {
args.card_id = std::stoi(require_value(argc, argv, index, "--card"));
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " --model yolo26n-det-int8-b1.tar.gz [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown argument: " + arg);
}
}
if (args.model.empty()) {
throw std::runtime_error("--model is required: yolo26n-det-int8-b1.tar.gz from the Model Zoo");
}
return args;
}

std::string shape_string(const std::vector<std::int64_t>& shape) {
std::string text = "[";
for (std::size_t index = 0; index < shape.size(); ++index) {
text += (index == 0 ? "" : ", ") + std::to_string(shape[index]);
}
return text + "]";
}

std::size_t element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const auto dim : shape) {
count *= static_cast<std::size_t>(dim);
}
return count;
}

// The MLA-only contract quantizes every tensor with one scale and one zero point:
// x = (q - zero_point) * scale
// q = clamp(round(x / scale) + zero_point, -128, 127)
const pcie::QuantParams& require_quant(const pcie::TensorInfo& info) {
if (!info.quant.has_value()) {
throw std::runtime_error("tensor '" + info.name + "' publishes no quantization parameters");
}
return *info.quant;
}

std::int8_t quantize(const float value, const pcie::QuantParams& quant) {
const float code = std::nearbyint(value / quant.scale) + static_cast<float>(quant.zero_point);
return static_cast<std::int8_t>(std::clamp(code, -128.0F, 127.0F));
}

float dequantize(const std::int8_t code, const pcie::QuantParams& quant) {
return static_cast<float>(static_cast<std::int32_t>(code) - quant.zero_point) * quant.scale;
}

const std::int8_t* int8_data(const pcie::Tensor& tensor) {
return reinterpret_cast<const std::int8_t*>(static_cast<const std::uint8_t*>(tensor.data) +
tensor.byte_offset);
}

// Preprocessing of the reference model: one RGB HWC image with pixels in [0, 1]. Resize, BGR to
// RGB, divide by 255, then quantize with the ingress parameters. Another model needs its own
// recipe.
std::vector<std::int8_t> quantize_image(const cv::Mat& bgr, const pcie::TensorInfo& input) {
if (input.shape.size() != 3 || input.shape[2] != 3) {
throw std::runtime_error("the reference model takes a three-channel HWC input, got " +
shape_string(input.shape));
}
const pcie::QuantParams& quant = require_quant(input);
cv::Mat rgb;
cv::resize(bgr, rgb,
cv::Size(static_cast<int>(input.shape[1]), static_cast<int>(input.shape[0])));
cv::cvtColor(rgb, rgb, cv::COLOR_BGR2RGB);
std::vector<std::int8_t> codes(rgb.total() * rgb.channels());
for (std::size_t index = 0; index < codes.size(); ++index) {
codes[index] = quantize(static_cast<float>(rgb.data[index] / 255.0), quant);
}
return codes;
}

} // namespace

int main(int argc, char** argv) {
try {
const Args args = parse_args(argc, argv);
if (!std::filesystem::is_regular_file(args.model)) {
throw std::runtime_error("model does not exist: " + args.model);
}
const cv::Mat image = cv::imread(kImagePath, cv::IMREAD_COLOR);
if (image.empty()) {
throw std::runtime_error(std::string("OpenCV could not decode: ") + kImagePath);
}
pcie::ConnectionOptions connection;
connection.card_id = args.card_id;

// CORE LOGIC
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

std::cout << "[OK] 027_run_mla_only_int8\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

소스​