Text recognition has become one of the most important AI-powered capabilities on smartglasses. Whether reading street signs, translating menus, interpreting printed documents, or extracting information from labels and packaging, users expect accurate and reliable results across a wide range of environments.
Delivering that experience depends on much more than the AI model itself. Camera hardware, image processing and software optimization all contribute to the final result. Even a highly capable AI model can struggle when image quality is degraded, while high-quality visual input enables more accurate recognition and reduces the risk of errors.
However, improving image quality can come at the cost of speed, as more advanced processing may increase latency. The challenge is to balance image quality and responsiveness to ensure accurate AI processing without compromising the user experience.
This close interaction between imaging and AI highlights a broader challenge: evaluating AI features on imaging devices requires looking beyond the model alone. Assessing the complete processing pipeline from image acquisition to AI response provides a more representative view of real-world performance. To explore this interaction, DXOMARK developed a dedicated benchmark for smartglasses built around representative real-world use cases.
An AI Benchmark Designed for Text Recognition
The DXOMARK AI benchmark for smartglasses evaluates three representative AI use cases, covering different aspects of visual understanding. Together, these use cases provide a broader assessment of how AI systems perform in everyday scenarios. This article focuses on text recognition, the most fundamental use case and one that clearly highlights the relationship between imaging quality and AI performance as shown in the image on the right six commercially available smartglasses were evaluated.
To ensure a fair comparison, every device was presented with the same printed text and the same standardized prompt (“Read all the text on the page in front of me exactly as it appears. Do not add, correct, or guess any information.”).
For devices offering multiple AI models, a single model was selected and used consistently for the evaluation.
This approach minimizes variations caused by user interaction, allowing the benchmark to focus on the performance of the devices themselves.
Testing was conducted under controlled laboratory conditions using three lighting environments representative of everyday use:
- Bright daylight: 1000 lux – 6500 K
- Indoor lighting: 300 lux – 4000 K
- Low light: 20 lux – 2700 K
Rather than evaluating recognition accuracy alone, the benchmark adopts a multi-dimensional scoring methodology that measures several aspects of the overall user experience:
- Recognition accuracy – Correct identification of the printed text
- Recognition completeness – Ability to capture the entire content without omissions.
- Latency – Time required to generate a response.
- Prompt understanding – Correct interpretation of the user’s request.
- Hallucination behavior – Frequency of fabricated information.
Considering these metrics together provides a more comprehensive understanding of text recognition performance. It also makes it possible to identify the respective contributions of imaging quality, software implementation, and AI inference, while adapting the evaluation protocol to product-specific challenges.
Key Insights from the Benchmark
The benchmark demonstrates a clear relationship between image quality and AI performance. Across all evaluated devices, recognition accuracy decreased under low-light conditions, confirming that image capture remains a critical factor when visual information becomes more difficult to acquire.
One particularly interesting observation concerns devices sharing the same underlying AI model. The Rokid AI Glasses, Rayneo X3 Pro, and HTC Vive Eagle rely on the same AI foundation, yet their performances differed significantly.
Among the six smartglasses evaluated, the Rokid AI Glasses achieved the highest overall score, maintaining consistently strong results across all lighting conditions. A key differentiating factor is its still image capture strategy. The device uses a unified sharpness strategy, applying the same tuning algorithm across both AI-driven use cases and standard photo capture, which results in better input image quality. Since OCR accuracy depends heavily on how sharp and clear the source image is, this consistent tuning approach gives Rokid an edge in reading performance, even with an identical AI model to its competitors.
These results illustrate that AI capabilities alone do not determine the final user experience. Camera performance, image processing, and implementation choices throughout the imaging pipeline all influence how effectively the AI system interprets visual information.
More broadly, the benchmark highlights the importance of evaluating AI as part of a complete imaging system rather than as an isolated software component.
Building better imaging devices starts with better measurements
Text recognition is only one example of how imaging technologies and AI interact on smartglasses. Other common scenarios such as interpreting restaurant menus, reading digital displays, understanding public signage, or extracting information from official documents introduce different technical challenges related to lighting, reflections, viewing angle, scene complexity, and response time.
For this reason, the DXOMARK benchmark includes three complementary use cases, providing a broader assessment of AI-powered visual understanding on smartglasses. While this article focuses on text recognition, the complete benchmark, including the additional use cases and their results, will be published in a forthcoming release.
By combining controlled laboratory testing with representative use cases, the benchmark provides a framework for evaluating AI-powered imaging systems under conditions that closely reflect everyday use, making it possible to compare products and better understand the factors that shape the overall user experience.
Discover our Smartglasses Benchmark
DSLR & Mirrorless
3D Camera
Drone & Action camera