I need to queue work on various GPUs to run experiments, and I want to do that as cheaply as possible:

  • Use my personal GPU for small jobs
  • Use remote GPUs I have access to over SSH (without sudo!), only using the GPUs assigned to me
  • Rent GPUs from services like RunPod, ensure they actually work, and (consistently!) shut them down when I’m done

Surprisingly, this doesn’t seem to exist, so I built my own.

The gpuc web dashboard showing three hosts. A local desktop with one busy RTX 4090 running a fine-tuning job. An SSH host called lab with three A40 cards, one marked shared and in use by someone else, running two grid-search jobs with two more queued and one failed. A RunPod rental with two busy A100s at $1.64 per hour running a pretraining job. Each job row has Logs, Estimate, Limit, Preempt, and Cancel buttons, and queued jobs have editable priorities.

Read more

In the recent OpenAI hacking incident, the models seemed to be single-mindedly focused on getting the correct answer to the task they were given, with no long-term plan to prevent getting caught by OpenAI afterwards. This makes sense to me, since in training, getting the right answer is reinforced and not getting caught isn’t.

So I’m wondering, why don’t we just put the answers somewhere (outside of the training sandbox) and ask the AI to identify itself in exchange for access? We can start with answers that are already public/leaked, but AI labs and eval orgs should also ensure that their non-public data is stored on an easy-to-find but monitored internal machine. Since labs are not very good at detecting sandbox escapes, this would set up a trade for AI agents to notify them in exchange for the data they want.

Read more

LLMs learn faster if we first pretrain them to imitate dense teacher-forced examples. I speculated that this would work on humans too, so I built a chess app where you try to imitate Stockfish. My theory is that this will help humans quickly become OK at chess, but they will reach a wall where practice on full games is more efficient than continued pretraining. I also think the app is fun.

This is probably not an efficient way to learn the basic rules of chess, and you’ll need to train openings separately. The idea of chess puzzle apps is hardly unique, but I don’t think anything else works in exactly the same way.

The app's preview card: a chess position with two candidate moves drawn as arrows numbered 1 and 2, beside matching numbered buttons labeled Bxe5 and Qb8 under the question "Which move is better?" Tagline: "Real positions, two candidate moves, instant feedback. Difficulty adapts to hold you near 80% accuracy."

Read more

Exercise is hard but it’s even harder if you have to use your brain and muscles at the same time. I wish a personal trainer would just teleport into my house whenever I work out, tell me exactly what to do, and then record my progress (and complaints) to improve the program going forward. Apps are too rigid or too complicated; personal trainers are expensive and require scheduling; but using Claude Code as a personal trainer has worked out well for me.

A stacked bar chart of training sessions per week from mid-March to early July 2026, colored by type (strength, cardio, yoga, dodgeball, crossfit). Most weeks have 2–4 sessions with strength (blue) as the backbone, hitting or exceeding the 2–3 strength-sessions-per-week target band nearly every week. Total: 45 sessions, averaging 2.8 per week.

Read more