Like many others, my first experience with writing code was creating a program that printed “Hello world” to the terminal. Thankfully, my teacher in 8th grade continued this tradition. Hello World is a poetic way to start learning. A new entity has emerged: programmer + computer. It’s so significant that the whole world must be addressed. The potential of this newly empowered person and machine is limitless.

The first lesson in programming is really about setting up the environment that lets you write and run code. This is unduly difficult for beginners. It’s quite a steep learning curve for somebody who doesn’t even know how to make a program that prints a line of text. Years before I had any formal instruction in programming, I yearned to write code to make computer games. I flailed and floundered because I couldn’t find answers on Google to my devastatingly simple questions. I didn’t know the right search terms. Now LLMs have demolished this problem. You can query humanity’s collective digital knowledge with questions in plain English. With LLMs, learning the basics of nearly any technical subject is suddenly accessible to anyone who asks.1

It’s the phenomena of LLMs that make me newly excited about robotics. Now a computer program can manipulate language with enough mastery that it is genuinely useful. Why can’t computers manipulate physical objects in the real world with enough mastery to be genuinely useful? The language used here is not English, but servo positions and camera and sensor inputs.

I suppose the answer is that right now there is not enough training data in the world of robotics. I don’t know how true this is. I’m a beginner in this field. I want to understand it better. That’s why I bought a robot arm.

I wanted the cheapest simplest piece of hardware that could still do something useful. My research led me to the SO-ARM101. This low-cost robot arm is teleoperated with another arm to create training data. This data is used to train a model to operate the arm. The arm began in 2024 as a collaboration between Hugging Face and Rob Knight.2 It was designed to use LeRobot, Hugging Face’s open-source library for teaching robots by example. Hugging Face the company is a decade old, and has evolved significantly. Its current form is often described as “the GitHub of machine learning.” Upon inception, its founders used the hugging face (🤗) emoji as a silly placeholder name. The name stuck, probably because of silliness. I include this info because I feel like everyone is perplexed by the name.

I bought my SO-ARM101 kit from Seeed Studio. At the time of writing, this is the best place to buy an all-in-one kit. Including the 3D-printed parts and two webcams, it cost me $353 total. This is shockingly affordable! The main driver of the price is the FeeTech ST3215 smart servos. They’re about $20 apiece and each arm needs six of them. I was surprised at how powerful these little servos are. I pinched my hand with the gripper, and it was so firm it was a little painful!

The setup instructions on the SO-ARM GitHub page are easy enough to follow. I had Codex CLI walk me through calibrating the motors, which made it a breeze. I then set up the environment for my first experiment: picking up a ball and placing it in a tray.

I teleoperated the robot to place the ball in the tray for the recommended 50 demonstrations. I used 5 different starting positions for the ball. A webcam captured video of each trial. It took only about 30 minutes to collect this dataset.

As suggested by the tutorial on Hugging Face, I trained the ACT policy on my dataset. This policy (aka model, algorithm, neural net) was introduced in the influential ALOHA paper. The major technical innovation introduced in this paper was a neural net that predicted not just the immediate next action of a robot’s motion, but a whole chunk of actions about two seconds long. This apparently improved performance.

I started training the policy on my data with my wimpy HP laptop. It was going to take 15 hours. I stopped this and submitted the training job to Hugging Face jobs. It took 20 minutes to train on a remote A100 GPU and cost me 30 cents! I am continually astounded at how affordable this project is, and how accessible the required knowledge is. What a time to be alive.

After training the policy, I deployed it to my robot arm. It shakily moved in an attempt to grab the ball, but always grasped too far to the left of the ball. Its jitters were intense. My robot had Parkinson’s. For this attempt I used training parameters suggested by Codex. I decided to give it another go with more thoroughly vetted parameters. I also used a camera mounted to the robot’s wrist rather than an overhead one.

After training for the second time, I witnessed magic. The robot moved quickly and grabbed the ball on the first try. I was surprised to see that while the policy was still running the robot would move the ball back into the tray if I took it out. It was significantly less jittery than before. I believe this must be normal because the SO-ARM gif on the LeRobot GitHub has the shakes worse than my robot.

I’ve copied the training parameters below.3 This is what worked for me, and maybe it will help someone else. I’ve also included the rollout parameters I used.4 Notably, I decreased the chunk size and n_action_steps from the default of 100, 100. Chunk size is how many future actions the model predicts, and n_action_steps is how many of those actions are actually executed before it predicts another chunk of actions. Curiously, it seems like most labs have had best success predicting more actions than the robot actually executes in each chunk. Also, I increased the robot.max_relative_target to 50 degrees. It defaults to 5. A low value prevents the robot from moving too fast. This setting is for safety, I suppose. It’s much too safe to be useful in my opinion. I’m tired of agonizingly slow robots!

It’s extremely gratifying to see this thing work. With only 50 examples, the model learned to control a robot without knowing anything about it beforehand! Next, I will deploy a more sophisticated model on this robot. ACT is three years old already, and I want to be closer to the cutting edge. Physical Intelligence has a foundation model that comes with thousands of hours of robot manipulation baked into the model, and should be able to perform a task without specific training.

My goal is to have the SO-ARM do something useful. Stay tuned!

lerobot-train.exe `

--dataset.revision=v3.0 `

--dataset.repo_id=USER/ball_tray_v1 `

--dataset.eval_split=0.1 `

--dataset.video_backend=pyav `

--policy.type=act `

--policy.repo_id=USER/act_ball_tray_v2 `

--policy.chunk_size=50 `

--policy.n_action_steps=25 `

--policy.vision_backbone=resnet18 `

--policy.pretrained_backbone_weights=ResNet18_Weights.IMAGENET1K_V1 `

--output_dir=outputs/train/act_ball_tray_v2 `

--job_name=act_ball_tray_v2 `

--seed=1000 `

--batch_size=16 `

--steps=10000 `

--eval_steps=500 `

--max_eval_samples=128 `

--save_freq=2000 `

--save_checkpoint_to_hub=true `

--wandb.enable=false `

--job.target=a10g-small `

--job.timeout=1hlerobot-rollout.exe `

--strategy.type=base `

--policy.path=USER/act_ball_tray_v2 `

--device=cpu `

--robot.type=so101_follower `

--robot.port=COM8 `

--robot.id=follower `

--robot.calibration_dir=.\calibration\robots\so_follower `

--robot.max_relative_target=50 `

--robot.cameras="{wrist: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" `

--task="Pick up the block and place it in the bin" `

--fps=20 `

--duration=60