Products — ArcleIntelligence

Our Product
suite

One model shipped, one architecture in research. Both built for absolute privacy and full on-device operation — no cloud, no compromise.

Fig. 3 — Achievement of Our Products

Benchmark
scores

Arcle V1 against published results for open models in the same parameter class. Reported in full — including where it trails.

Swipe or drag to explore

77.5
38.4
75.6
80.6
88.6
Arcle V1
5.84B
Gemma 3
4B
Llama 3.2
3B-It
Qwen2.5
3B-It
Phi-4-mini
3.8B
GSM8K
multi-step mathematical reasoning
80.0
82.4
n/a
n/a
n/a
Arcle V1
5.84B
Gemma 3
4B
Llama 3.2
3B-It
Qwen2.5
3B-It
Phi-4-mini
3.8B
ARC-Easy
grade-school science reasoning
67.0
77.2
77.2
74.6
69.1
Arcle V1
5.84B
Gemma 3
4B
Llama 3.2
3B-It
Qwen2.5
3B-It
Phi-4-mini
3.8B
HellaSwag
commonsense sentence completion
48.5
56.2
76.1
82.6
83.7
Arcle V1
5.84B
Gemma 3
4B
Llama 3.2
3B-It
Qwen2.5
3B-It
Phi-4-mini
3.8B
ARC-Challenge
hard science reasoning
43.5
59.6
61.8
65.0
67.3
Arcle V1
5.84B
Gemma 3
4B
Llama 3.2
3B-It
Qwen2.5
3B-It
Phi-4-mini
3.8B
MMLU
broad multi-domain knowledge

Methodology — Arcle V1 figures are measured through Arcle's own evaluation harness at n = 400 samples per benchmark, after the reasoning-SFT stage. Peer figures are published results: Gemma 3 4B from the Gemma 3 technical report; Llama 3.2 3B-Instruct, Qwen2.5 3B-Instruct and Phi-4-mini from the Phi-4-mini-instruct model card. Because harnesses, sample counts and shot counts differ between sources, these comparisons are indicative rather than like-for-like. Blank bars mark benchmarks with no published figure from that source.