#direct manipulation

Direct manipulation (in the GUI sense) / 7 posts

Newsletter
  • Get a newsletter digest every Friday:
Contact

Send me feedback, tips, bug reports, and anything else:

About me and the blog

My name is Marcin Wichary. I’ve worked as a UX designer, typographer, front-end person, and manager at Google, Medium, and Figma. I’ve also written a book about keyboards and typing, gave talks, and published essays about design, typography, and technology.

Unsung is my blog about software craft and quality.

More info about Unsung

Send me feedback, tips, bug reports, and anything else:

More info about Unsung

“Act now, apologize later”

Fighting games require rather precise timing: the action is frenetic, the inputs rapid, and the outcomes can be determined by what players do down to – sometimes – individual one-sixtieth-of-a-second frames.

This is fine when you play locally, but gets infinitely more tricky the moment you attempt to kick someone’s more remote ass:

Information sent to your opponent may be delayed, arrive out of order, or become lost entirely depending on dozens of factors, including the physical distance to your opponent, if you’re on a WiFi connection, and whether your roommate is watching Netflix.

The part of the game responsible for dealing with networking is called netcode, and the classic solution to the lag problem has been delay-based netcode with a sort of lowest common denominator approach: if your opponent has a “ping” of 50ms, your actions will also be delayed by 50ms, to ensure fairness.

It’s not that hard to imagine many problems with this idea. Lag comes and goes but the game doesn’t smooth it out, the play feels like jello with even an okay connection, and the game occasionally freezes, waiting for a suddenly dropped connection to resume, not knowing what to do in absence of information.

That’s why within the last decade, what gained popularity is a new kind of netcode called rollback netcode, which operates by… lightly predicting the future. Most of the time, within a reasonable lag, the game will feel as fast as a local game to both players – and even when the optimistic prediction doesn’t match what eventually happens in reality and the code has to roll it back, visibly, in front of a player, the change might be too small to notice and only rarely it’ll result in rubber banding.

Core-A Gaming has a great video about specifics of rollback netcode:

It covers a lot of stuff, including nuances like drift and design decisions that game makers might want to incorporate to soften rollback netcode’s most annoying side effects. (Also, the guitar playing bit at 0:22 is absolutely brilliant.)

The video is largely based on an earlier essay by Infil, which has even more details (audio!) and which I excerpted above.


It’s not just videogames. When Apple announced in 2019 that they reduced the latency when using their iPad Pencil from 20ms to 9ms, they didn’t hide that they had to peek into the future as well:

The specifics were, as far as I know, never revealed. But the fact that some of the input events the app receives are not actually real is official and documented:

It takes time for UIKit to generate and deliver touch events to your app, and it takes time for your app to process those events and render the results. In fact, it takes enough time that there can be a visible lag between the movements of a person’s finger or Apple Pencil and the rendered results. To minimize the perceived latency between touch input and rendered content, you can incorporate predicted touches into your event handling.

In fighting games, the most accurate prediction is a trivial “whatever buttons and controls the player holds on the current frame are likely the same they will hold on the next frame,” but for other games – and for Apple’s Pencil – slightly more advanced prediction might be necessary:


It feels disappointing that no matter what we do, there might be no way to reduce latency below a certain threshold without cheating.

But if there’s any consolation, it’s that cheating seems to be the only answer, and that has been the case long before computers came aboard. There are many documented ways our brains have their own netcode to cover for various lags occurring within our bodies. There’s the stopped-clock illusion, the cutaneous rabbit effect, the color phi phenomenon, and – most relevant for today, even though it has the most boring name – the flash-lag effect, documented on this nice page by Michael Bach:

The flashing lines are shown exactly at the same time and angle as the rotating line, but it certainly doesn’t appear to be the case. Just like fighting games and Apple Pencil do, our brains predict and extrapolate the movement of the moving line so that our other senses can intercept it at the right moment. The go-to example of why it’s needed is that you wouldn’t otherwise ever be able to catch a ball thrown your way – but I imagine it might be useful in a street fight, as well.

For a lotmore technical aspects of rollback netcode, check out this long conference presentation from GDC.

Unsung Heroes: Super Sprint

I know, I know. I’m supposed to say iPod’s click wheel, or the Western Electric 500 rotary dial, or maybe the first Nest.

But, have you ever played Super Sprint?

It was Super Sprint that had the first amazing rotary controller I’ve ever used, and the whole cabinet design told you the game was very well aware of that.

Super Sprint was a 1986 arcade game from Atari that was, in a nutshell: four cars (at least one computer-controlled), eight tracks, fast races.

You might think that those wheels functioned similarly to a regular car steering wheel, but not really – they were much easier to spin, and needed to travel further to rotate the car. You wouldn’t steer delicately, but rather the opposite: you needed to throw the wheel violently in one direction, and then, at the perfect moment, stop it on a dime:

So the huge wheels were not realistic. Neither were the cars. They accelerated rapidly – the gas pedal was your only other control – and had a ridiculous amount of understeer.

Don’t let the size and intensity of the interface fool you, though: this was a very precise operation. The game was tight – Rollercoaster Tycoon tight or Excel 97 tight, a whole decade before them. In the world awash with slow computers, Super Sprint lived up to its name, laughing latency and delays in their faces. It’s hard for me, even today, to imagine something faster or tighter – and for even my contemporary work, it’s good to remember things can feel this way. And, on top of all that? Better-than-usual sound design, higher-than-usual resolution, shortcuts to reward really good players, and a bunch of great details and easter eggs.

Super Sprint wasn’t Atari’s first attempt here – it was preceded by Sprint 2, Sprint 4, Sprint 8, and Sprint One – and you could tell. It was designed and coded by Kelly Turner and Robert Weatherby, polished as hell, and might have been the first interface between the person, the hardware, and the software that really inspired me. It was almost as much fun to watch three good players compete, as it was to play yourself. But when I played it, it’s possible these were my first – please excuse me here – motor memories.

Atari used the same wheel for other games, famous and obscure, but this was where it met its match in software. I can show that to you, but experiencing it is impossible from afar. Even perfect emulation can’t do it justice – there has simply never been a home controller that approached it. (I mean, the whole cabinet weighed 400 pounds!)

But if you are ever in an old-school-themed arcade, look out for Super Sprint (locations), or its two-player cousin Championship Sprint (locations) – and give it a turn or two.

Mark MacKay’s explainers

If you haven’t played Kern Type yet, you’re in for a treat – released by Mark MacKay in 2011, it’s a delightful browser game that teaches you the art and the math of text kerning:

If you have played it, you’re still in for a treat, as MacKay under the Method of Action label released a few similar games since:

They’re interesting not just because you can explore the feel of visual, typographical, and color craft by play – but also because they’re also really nicely made explainers. Here, you can learn things, but also learn about making things that teach things.

For example, I liked the thoughtful onboarding, animations, and sound design of The Boolean Game, the keyboard navigation in Kern Type, and this general idea present in a few games that it’s fun to learn by fixing slightly broken things, rather than starting from scratch.

As far as I can tell, MacKay also occasionally revisits the older games, so Kern Type today might feel much better than when you played it last. MacKay also vibecoded a quick aspect ratio game, and is thinking of a new game to teach text editing, which excites me to no end.

I added a new tag to Unsung, #explainer, to cover explainers and playgrounds like these.

Five moments in snapping history

Bear (a notetaking tool) has simple image resizing, with one extra nicety – if your images are near each other, resizing one will snap to the width of the other:

In the Finder, columns snap to the width necessary to keep all the names untruncated – and not one pixel more:

(Sidebar: This is also the only place I’m mentioning today that nicely uses the trackpad’s haptic feedback at the snap moment. I tried to indicate it in the video; this is not the final visual treatment I’m thinking of, but let me know if this kind of visualization of haptics feels useful to you!)

macOS does something really interesting when you get its windows close to each other. Instead of typical snapping – pulling the thing you hold toward the other item like a magnet – it instead prevents you from going further for a while, in either direction. Perhaps the right analog here would be glue:

If initially feels a bit funny, but I think I like it. It’s less aggressive and avoids needing some sort of cancellation (an option or a modifier key) if you don’t want it, because it never feels in a way.

It also works for matching heights, like in Bear:

In Figma, building atop regular snapping, we introduced something I awkwardly called “self-snapping”: if your objects are inside a container, the container will snap its padding to whatever it sees on the other side, without any explicit auto layout/​flexbox:

I am sharing these five examples (and one from before) because I think they exemplify a nice thing: precision without bureaucracy. In each case you could imagine an explicit heavy option somewhere in the menu…

  • Bear: Set Image Width…
  • macOS: Match Window Heights
  • Finder: Restore Column Width
  • Figma: Unify Padding

…that would feel slow and cumbersome.

Instead, these take the freedom of direct manipulation and sprinkle just enough almost-invisible structure in a moment where that structure is undeniably useful. (Of course, you still might want explicit options somewhere in the menu or your command palette, if only for accessibility reasons.)

I like that these quiet features have your back and make you look good, and that their creators understand that something almost aligned can feel worse than something completely misaligned.

The 1990s called and they want their dialog box back

This is perhaps my favourite feature in Lightroom. You press ⇧T, you draw a few lines, and presto – your photo is now even:

This is doubly magical to me. The first part is that this is even possible – that you can straighten the photo in both dimensions after the fact, and save for some parallax nuances the viewer won’t know any better.

For decades, this has been the domain of tilt-shift lenses, but if you ever tried to use one, you know how harrowing of an exercise this is. A tilt-shift lens looks more like a medical device and less like a piece of photography equipment:

The “obvious” way to emulate a tilt-shift lens in software is a bunch of sliders, and Lightroom has those also…

…but that’s still pretty cumbersome in practice, abstracted in a strange ways, like piloting a plane by pulling the linkages connected the flying surfaces: you will admire someone who can do that, but won’t ever want to do it yourself.

Hence the second magical moment: The team created the new interface I showed at the beginning, where you point to things that should be straight directly, and the necessary tilt-shift calculations happen behind the scenes.

Alas, Lightroom didn’t fully stick the landing. The interface is a bit jittery, and missing nice transitions that could help understand what’s going on. But what brought me here was this unpleasant interaction:

What’s wrong with it? If you want to play along, stop here and ponder: How would you improve it? Because this is a classic UI exercise where there are symptoms, and there are problems, and there are principles under the hood of it all.

The first possible improvement: Don’t do a dialog like this. These are ancient and so annoying. Every time I see a centered dialog covering everything, popping up in response to a delicate mouse operation, I want to shout “read the room!” It’s better to drop a little tooltip next to the cursor that automatically disappears: more modern, and more “compatible” with mousing.

Then: Why am I allowed to start and finish an action that the machine already knows won’t go anywhere? Disable the drawing option, put a little “verboten” icon on the mouse pointer, or do something else that will prevent me from drawing a line to begin with.

But that brings us to point three, and how I would approach this as a designer. Because I would – counterintuitively – go the other way and allow the user to draw as many lines as they wanted, and just didn’t permit to commit the entire operation if there were more than four lines on the screen.

Why is that?

It’s the same principle as you see in all the social media composing fields, and in well-trained forms: do not constrain the editing process.

This field is limited to 300 characters, but it’s clever enough to only enforce its limits when you try to post. There is no downside to allowing you more room in the editing process. Maybe you write by constructing a few sentences first and only then combining them into one, maybe you want to see two riffs one below the other to choose the better one, or maybe – this is most likely – you’re not even paying attention and your motor memory is doing the editing for you, instinctively. Use any text editor for just a few months, and cut, copy, and paste, word swapping, and splitting sentences become second-nature gestures – that is, until the UI starts throwing in some arbitrary barriers.

Above in Lightroom, it might actually be easier for me to draw a fifth line and then delete a previous one, instead of doing it in the precise order Lightroom desires, or by dragging an existing line to move it instead of creating a new one.

Maybe an overarching principle would be this: If you are aiming to build something so delightfully direct manipulation as Lightroom did here, you have to fully commit to that stance, even deep in the weeds. Because every time I see a 1990s dialog appear when my fingers are flying fast, I feel like this:

And something tells me others will too.

Tales of direct manipulation, pt. 1

Mac allows you to assign keyboard shorcuts to menu items, but the interface is clunky – you have to select the app even if you just came from it, and then type in the menu item name by hand without any assistance:

Other tools, like Keyboard Maestro, do something similar. You either have to type it again, or you can point to it, but in a replica of the menu of the app shown in a very different style and orientation:

But this week I learned of another app, KeyCue, that approaches this differently. You simply point to the menu item and hold the desired key for a while:

Okay, this is not a universal endorsement. The feature works clunkily, and KeyCue as a whole is way too comfortable adding itself to login items without asking.

But as far as singular interactions go, this is great and eye-opening. It made me realize that the previous things I’ve shown – System Settings, Keyboard Maestro – are really not GUIs, and they don’t practice direct manipulation. They’re still partially command line interfaces dressed up in GUI clothing.

We kind of lightly made fun of Jony Ive going angelic on “staying true to the material” and things being “beautifully, unapologetically plastic.” And there is, of course, value in command line and those kinds of approaches. But this part of KeyCue at least is unapologetically a graphical user interface, and it is nice to still be surprised in this space.

A nice moment in screenshotting on iOS

In iOS, I like how cropping quietly snaps to things that look like borders, with gentle haptics, without announcing anything: