- EN
- CN
A smart desktop companion robot built for toys, smart speakers, smart control hubs, and any voice-interaction product that needs the power of large language models. Powered by the ESP32-S3-WROOM-1 module with a 1.85" QSPI round touchscreen and a dual-microphone array, it supports offline wake-word detection and sound-source localization. Combined with LLM capabilities, it delivers two-way voice interaction, multimodal recognition, and agent control — a solid foundation for developers building complete on-device AI experiences.
Proactively detects user intent and shifting emotions. Drawing on the semantic understanding of large language models, it interprets tone, word meaning, and conversational context together, then responds with lifelike dynamic expressions and voice feedback — enhancing emotional expression and giving the device real personality.
Continuously records multi-turn conversations, remembering the user's name, preferences, and frequently used phrases, then recalling them in later interactions. The result is a personalized experience that adapts to each user's habits and raises the device's value as a true "emotional companion."
Pairing a motor-control module with the dual-microphone array, the robot tracks direction accurately across a 180° range. Every time you call out, it identifies the sound source and rotates its base to make eye contact — making every interaction more immersive and natural.
Switch tones and styles freely to build your own signature AI voice identity. With DIY voice persona support, the voice experience is more open and more personal.
Supports the MCP protocol and Function Call capabilities to connect with local smart devices — enabling remote control, task dispatch, and status feedback. It serves as a local control hub within a smart-home system, offering stable, efficient edge control and open expansion interfaces.
Deeply integrated with leading LLM platforms, including Amazon Nova, OpenAI, XiaoZhi AI, and Gemini, giving developers flexible choices.
Combines voice and touchscreen interaction, and senses device posture changes through an IMU sensor for richer, more varied ways to engage.
Videostrong is a leading provider of AI robotics hardware and end-to-end solutions. With 14 years of OEM/ODM experience, products and services in 60+ countries and regions, and nearly 100 million households served, we deliver full-chain solutions — from product design and structural development to software customization and mass-production delivery.
ESP32-S3-WROOM-1-N16R16VA (16MB Flash + 16MB PSRAM)
2.4 GHz Wi-Fi and Bluetooth 5 (LE)
1.85" QSPI round touchscreen, 360×360 resolution
Built-in 3W mono speaker; dual LMA3729T381-OY3S microphone array with offline wake-word detection and sound-source localization
microSD card slot (up to 32GB)
BMI270 6-axis IMU for posture sensing
USB-C (power/download), magnetic connector (expansion), Pogo pin
5V DC and 3.7V lithium battery (typical 700mAh), USB-C charging
Full-duplex voice interaction, multimodal recognition, LLM application support, rotating-base linkage
Discuss your OEM/ODM requirements with our engineering team and get a customized AI solution.
Contact Expert
Copyright © 2011-2025 Videostrong Technology Co., Ltd. All Rights Reserved 粤ICP备17154177号