MoLE - Modalix Language Model Evaluator
Overview
MoLE (Modalix Language Model Evaluator) is a benchmarking tool for evaluating the accuracy and performance of LLMs running on the Modalix platform.
It extends EleutherAI's lm-evaluation-harness and supports two backends:
- hf — runs evaluation on the host using HuggingFace transformers (baseline reference)
- modalix — runs evaluation on a Modalix board via the
llima benchmark-server
Installation
MoLE is a host-side benchmarking tool. Install and run it on the host machine outside the SDK Docker container, not from the SDK container and not on the Modalix device. The Modalix device only needs the LLiMa runtime and the llima benchmark-server process. See Neat Framework installation for the runtime installation flow.
Install MoLE on the host using sima-cli:
host:~$ sima-cli neat install llima/mole
This installs MoLE into a host virtual environment at ~/sima-mole-venv.
Usage
First, activate the MoLE virtual environment:
host:~$ source ~/sima-mole-venv/bin/activate
MoLE is then invoked via the llima-benchmark CLI with two subcommands. The <model_id> argument is always the HuggingFace model ID (e.g., meta-llama/Llama-3.2-3B-Instruct). In -b modalix mode this is not just a display label: it must match the tokenizer and config used to compile the deployed board model, because the board returns token scores only and does not provide tokenizer metadata.
Accuracy Benchmarking
Evaluates model quality against standard tasks:
(sima-mole-venv) host:~$ llima-benchmark accuracy <model_id> -b modalix \
-t <task> \
-o <output_dir> \
--max_num_tokens <max_num_tokens> \
--board_ip <board_ip> \
--board_model <model_path_on_board>
| Argument | Description |
|---|---|
model_id | HuggingFace model ID (e.g., meta-llama/Llama-3.2-3B-Instruct). For -b modalix, this must match the deployed model's tokenizer/config. |
-b | Backend to use: modalix (run on board) or hf (run on host as reference baseline). |
-t | Required. One or more evaluation tasks. Example tasks: hellaswag, triviaqa, piqa, winogrande, wikitext. See the task list for all available tasks. |
-o | Output directory for benchmark results. |
--board_ip | IP address of the Modalix board. Required for -b modalix. |
--board_model | Path to the compiled model directory on the Modalix device (e.g., /media/nvme/llima/models/Llama-3.2-3B-Instruct-a16w4). Required for -b modalix. |
--max_num_tokens | Maximum context length. Must be equal to or smaller than the value used during compilation. |
-n, --num_samples | Number of samples to evaluate. Runs the full task set if not specified. |
--board_ssh_user | SSH username for the Modalix board. Optional, default: sima. # |
--board_ssh_pass | SSH password for the Modalix board. Optional. Set to enable non-interactive automated benchmarking. |
Accuracy and loglikelihood benchmarking with -b modalix requires the deployed model to be compiled with --return_logits. This flag is off by default. See Model Compilation. If the model was compiled without this flag, the benchmark fails with: model not compiled with --return_logits; accuracy/loglikelihood tasks are unsupported.
In -b modalix mode, result tables are labeled as Modalix backend results and include the board target. The HuggingFace model_id still appears because MoLE uses it for tokenization and task metadata.
To use the HuggingFace backend as a reference baseline:
(sima-mole-venv) host:~$ llima-benchmark accuracy <model_id> -b hf -t <task> -o <output_dir>
For all available options, run llima-benchmark accuracy -h.
Performance Benchmarking
Measures Time To First Token (TTFT) and Tokens Per Second (TPS) on a Modalix board for different input lengths:
(sima-mole-venv) host:~$ llima-benchmark perf <model_id> \
-o <output_dir> \
--board_ip <board_ip> \
--board_model <model_path_on_board> \
--max_num_tokens <max_num_tokens> --max_new_tokens <max_new_tokens> \
--input_lengths 1024 2048 3072 4096
| Argument | Description |
|---|---|
model_id | HuggingFace model ID (e.g., meta-llama/Llama-3.2-3B-Instruct). For Modalix performance runs, this should match the tokenizer/config for the deployed model. |
-o | Output directory for benchmark results. |
--board_ip | IP address of the Modalix board. |
--board_model | Path to the compiled model directory on the Modalix device (e.g., /media/nvme/llima/models/Llama-3.2-3B-Instruct-a16w4). |
--max_num_tokens | Maximum context length. Must be equal to or smaller than the value used during compilation. |
--max_new_tokens | Maximum number of tokens to generate in the output. |
--input_lengths | Optional exact input-token lengths to benchmark. Values must be unique and each value plus --max_new_tokens must fit within --max_num_tokens. When omitted, MoLE generates automatic power-of-two buckets. |
--board_ssh_user | SSH username for the Modalix board. Optional, default: sima. |
--board_ssh_pass | SSH password for the Modalix board. Optional. Set to enable non-interactive automated benchmarking. |
For all available options, run llima-benchmark perf -h.