Skip to main content

Run the MLA Only with INT8 Tensors

Run the MLA Only with INT8 Tensors — animated walkthrough overview

FieldValue
CategoryPCIe Co-Processing
DifficultyIntermediate
Estimated Read Time15 minutes
LabelsPCIe, MLA, INT8, quantization, tensor

One program quantizes an image on the host, runs the MLA-only route, dequantizes the heads, and checks them against the default route on the same queue.

Walkthrough​

Inspect the MLA-only contract​

Construct the Model with mla_only enabled. info() now reports INT8 inputs and outputs, each with quant.scale and quant.zero_point. The reference model has one input, images, an INT8 [640, 640, 3] HWC tensor.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

Quantize on the host​

The reference model expects one RGB image with pixels in [0, 1]. Resize the image to 640x640, convert BGR to RGB, divide by 255, and apply the quantization equation with the ingress parameters. This preprocessing belongs to the model, not to the archive contract: another model needs its own recipe. Keep the dequantized values of the same codes: they are the exact FP32 input the default route needs for a like-for-like comparison.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

Run the INT8 route​

build() starts a card pipeline that contains only the MLA. run() accepts the INT8 tensors and returns one dense INT8 tensor per output, in the order and with the names that info().outputs reported. The route rejects any other dtype; an FP32 push fails instead of being quantized on the card.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

Dequantize and compare​

Build a second Model without mla_only and send the dequantized FP32 values through the default route. Dequantize the INT8 heads with each output's parameters and print the largest deviation per head in units of that head's scale. Both routes execute the same MLA program on the same codes, so the error is zero.

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

Run​

Install the PCIe host package and download the tutorial bundle as described in Tutorial Setup.

The tutorial runs the Model Zoo YOLO26n INT8 archive, which is compiled for direct MLA input and output. Download it into the extracted PCIe extras root:

sima-cli download https://docs.sima.ai/pkg_downloads/SDK2.1.3/models/modalix/yolo26-detection/yolo26n-det-int8-b1.tar.gz

Other archives qualify when they were compiled with the Model SDK tessellate_parameters enable_mla=True, an HWC DRAM layout on every input, and HWC16 on every output. Archives that tessellate on the CVU instead, such as the Model Zoo yolo_v8s build, are rejected when mla_only is enabled:

mla_only does not support stage 'tessellate_quantize_0_MLA_0/...' (tess)

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

C++ (build from source):

./build.sh --target tutorial_027_run_mla_only_int8
./build/tutorials-standalone/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

Both versions print the contract and a zero deviation for every head:

MLA-only contract:
input images INT8 [640, 640, 3] scale=0.00390434 zero_point=-128
output bbox_0 INT8 [80, 80, 4] scale=0.0302856 zero_point=-117
...
Dequantized MLA-only outputs vs the default route (error in scale units):
bbox_0 [80, 80, 4] max_err=0.0000
...
[OK] 027_run_mla_only_int8

The default is card 0 and queue 0. Pass --card N only when using another card.

In Practice​

For an INT8 archive, enable mla_only when the application owns quantization: it already produces INT8 from a sensor or an earlier model, it needs the raw INT8 heads for its own postprocessing, or it wants to remove the quantize and dequantize stages from the card-side latency. Read every scale and zero point from model.info(); never copy them from another build of the model.

The route is all or nothing. Every input must match the dtype, shape and byte size of its info().inputs entry, and image preprocessing or box decode cannot be combined with mla_only. Keep the default route when the host holds FP32 data and does not need to own the conversion.

For deployment diagnostics, continue with the PCIe model workflow.

Full source​

Show the complete source programs
pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
// Quantize on the host, run only the MLA over PCIe, and dequantize the INT8 results.
//
// Usage:
// tutorial_027_run_mla_only_int8 --model yolo26n-det-int8-b1.tar.gz [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <opencv2/imgcodecs.hpp>
#include <opencv2/imgproc.hpp>

#include <algorithm>
#include <cmath>
#include <cstdint>
#include <cstdlib>
#include <filesystem>
#include <iomanip>
#include <iostream>
#include <stdexcept>
#include <string>
#include <utility>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kRunTimeoutMs = 30000;
constexpr double kMaxErrorScales = 0.05; // one twentieth of a quantization step
constexpr char kImagePath[] = "share/sima-pcie-host/tutorials/assets/street-scene.png";

struct Args {
std::string model;
int card_id = 0;
};

std::string require_value(int argc, char** argv, int& index, const char* option) {
if (index + 1 >= argc) {
throw std::runtime_error(std::string("missing value for ") + option);
}
return argv[++index];
}

Args parse_args(int argc, char** argv) {
Args args;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--model") {
args.model = require_value(argc, argv, index, "--model");
} else if (arg == "--card") {
args.card_id = std::stoi(require_value(argc, argv, index, "--card"));
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " --model yolo26n-det-int8-b1.tar.gz [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown argument: " + arg);
}
}
if (args.model.empty()) {
throw std::runtime_error("--model is required: yolo26n-det-int8-b1.tar.gz from the Model Zoo");
}
return args;
}

std::string shape_string(const std::vector<std::int64_t>& shape) {
std::string text = "[";
for (std::size_t index = 0; index < shape.size(); ++index) {
text += (index == 0 ? "" : ", ") + std::to_string(shape[index]);
}
return text + "]";
}

std::size_t element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const auto dim : shape) {
count *= static_cast<std::size_t>(dim);
}
return count;
}

// The MLA-only contract quantizes every tensor with one scale and one zero point:
// x = (q - zero_point) * scale
// q = clamp(round(x / scale) + zero_point, -128, 127)
const pcie::QuantParams& require_quant(const pcie::TensorInfo& info) {
if (!info.quant.has_value()) {
throw std::runtime_error("tensor '" + info.name + "' publishes no quantization parameters");
}
return *info.quant;
}

std::int8_t quantize(const float value, const pcie::QuantParams& quant) {
const float code = std::nearbyint(value / quant.scale) + static_cast<float>(quant.zero_point);
return static_cast<std::int8_t>(std::clamp(code, -128.0F, 127.0F));
}

float dequantize(const std::int8_t code, const pcie::QuantParams& quant) {
return static_cast<float>(static_cast<std::int32_t>(code) - quant.zero_point) * quant.scale;
}

const std::int8_t* int8_data(const pcie::Tensor& tensor) {
return reinterpret_cast<const std::int8_t*>(static_cast<const std::uint8_t*>(tensor.data) +
tensor.byte_offset);
}

// Preprocessing of the reference model: one RGB HWC image with pixels in [0, 1]. Resize, BGR to
// RGB, divide by 255, then quantize with the ingress parameters. Another model needs its own
// recipe.
std::vector<std::int8_t> quantize_image(const cv::Mat& bgr, const pcie::TensorInfo& input) {
if (input.shape.size() != 3 || input.shape[2] != 3) {
throw std::runtime_error("the reference model takes a three-channel HWC input, got " +
shape_string(input.shape));
}
const pcie::QuantParams& quant = require_quant(input);
cv::Mat rgb;
cv::resize(bgr, rgb,
cv::Size(static_cast<int>(input.shape[1]), static_cast<int>(input.shape[0])));
cv::cvtColor(rgb, rgb, cv::COLOR_BGR2RGB);
std::vector<std::int8_t> codes(rgb.total() * rgb.channels());
for (std::size_t index = 0; index < codes.size(); ++index) {
codes[index] = quantize(static_cast<float>(rgb.data[index] / 255.0), quant);
}
return codes;
}

} // namespace

int main(int argc, char** argv) {
try {
const Args args = parse_args(argc, argv);
if (!std::filesystem::is_regular_file(args.model)) {
throw std::runtime_error("model does not exist: " + args.model);
}
const cv::Mat image = cv::imread(kImagePath, cv::IMREAD_COLOR);
if (image.empty()) {
throw std::runtime_error(std::string("OpenCV could not decode: ") + kImagePath);
}
pcie::ConnectionOptions connection;
connection.card_id = args.card_id;

// CORE LOGIC
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

std::cout << "[OK] 027_run_mla_only_int8\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

Source​