Flow as defined by Mihály Csíkszentmihályi, and casually: designing experiences that allow you to do things without thinking about the interface / 62 posts
In response to a previous post, one of the readers wrote this:
I would phrase this as the fault of the animation having violated the Hippocratic oath for animations, which is “first, do no harm” a.k.a. “first, don’t interfere with user input”.
Yes, one has to be afraid of hyperbole – was there ever an onscreen transition that saved a life? – but there is something I really liked about this phrasing. After all, most transitions and animations are decorators, and every transition and animation is, by definition, also a delay.
Here’s Shazam as I open it, and overlaid are my frantic taps as I’m trying to make it start the song recognition process:
It’s a cute cold start animation, but:
It’s not interruptible, like animations should be.
It doesn’t even buffer the taps, so – in case the animation covers for something truly uninterruptible, like loading from the cloud – I can’t simply tap and forget, but instead have to wait and tap after it’s done animating.
What’s particularly frustrating here is how Shazam is being used. On the other end of this whole flow is a song that might already be fading out – timing myself, a user, can’t control – so time wasted on the uninterruptible animation might be seconds separating failure from success.
I know it does feel like a blink of an eye when watched out of context, and it is literally just a bit over a second of a delay in the best case scenario, but I wanted to share it as a general example. I believe these seconds add up, especially in a stressful moment, on an older device, repeated many times a day, across different apps… or all of the above. (My go-to example: Imagine your keyboard keys reacting with a second of delay!)
At least, I should compliment Shazam for not making another mistake: after the recognition starts, tapping the same button doesn’t cancel it – instead, you have to tap a close box in the corner. Here, the designers realized I might actually be slamming the big button many times over, and I shouldn’t be punished for it even more.
Buttondown, a newsletter publisher, has a pretty standard CSS editor in its web app. You can edit the code and whenever you make any change, you can then press the Save button to make it go live:
The view also thoughtfully supports pressing ⌘S to do the same thing:
However, if there is nothing to save, ⌘S is ignored by the CSS editor, and falls back to the browser’s handling:
It feels logical: ⌘S only takes effect when the button is visible, otherwise why would a user press it? In a front-end sense, it might even seem thoughtful. We all witnessed web apps that greedily took over some interaction, and broke things in the process. Hell, I did that myself.
In theory, you – the user – take a careful look at the state of things, notice the button, and then press ⌘S.
But in practice? It’s none of the above. Fingers don’t look around. You might press ⌘S twice in a row. Or after an undo to an already saved state. Or after pressing another key so light it didn’t register. Or after just sitting down to an open document that’s already saved. Or because you weren’t sure if the previous ⌘S press worked. Or just to be sure. You might not know why, and you might not even notice. The beautiful raw power of ⌘S as a citizen of motor memory is that it’s automatic, mindless, habitual.
That’s why ⌘S here needs to be both deterministic and idempotent – if there is nothing to save, don’t let it fall back to the browser, don’t show an error message, don’t beep at the user. Just ignore the keystroke altogether.
It might feel funny, but there are tons of places in the UI that already look the other way. Off the top of my head:
if you center align already center-aligned text in any writing app, the app just ignores you,
if you press ⌘A to select all more than once, no one’s shouting at you,
if you try to click a disabled button, the click gets quietly swallowed.
It gets a lot more interesting than these, but we’ll talk about more examples of “finger logic” in future posts.
I know, I know. I’m supposed to say iPod’s click wheel, or the Western Electric 500 rotary dial, or maybe the first Nest.
But, have you ever played Super Sprint?
It was Super Sprint that had the first amazing rotary controller I’ve ever used, and the whole cabinet design told you the game was very well aware of that.
Super Sprint was a 1986 arcade game from Atari that was, in a nutshell: four cars (at least one computer-controlled), eight tracks, fast races.
You might think that those wheels functioned similarly to a regular car steering wheel, but not really – they were much easier to spin, and needed to travel further to rotate the car. You wouldn’t steer delicately, but rather the opposite: you needed to throw the wheel violently in one direction, and then, at the perfect moment, stop it on a dime:
So the huge wheels were not realistic. Neither were the cars. They accelerated rapidly – the gas pedal was your only other control – and had a ridiculous amount of understeer.
Don’t let the size and intensity of the interface fool you, though: this was a very precise operation. The game was tight – Rollercoaster Tycoon tight or Excel 97 tight, a whole decade before them. In the world awash with slow computers, Super Sprint lived up to its name, laughing latency and delays in their faces. It’s hard for me, even today, to imagine something faster or tighter – and for even my contemporary work, it’s good to remember things can feel this way. And, on top of all that? Better-than-usual sound design, higher-than-usual resolution, shortcuts to reward really good players, and a bunch of great details and easter eggs.
Super Sprint wasn’t Atari’s first attempt here – it was preceded by Sprint 2, Sprint 4, Sprint 8, and Sprint One – and you could tell. It was designed and coded by Kelly Turner and Robert Weatherby, polished as hell, and might have been the first interface between the person, the hardware, and the software that really inspired me. It was almost as much fun to watch three good players compete, as it was to play yourself. But when I played it, it’s possible these were my first – please excuse me here – motor memories.
Atari used the same wheel for other games, famous and obscure, but this was where it met its match in software. I can show that to you, but experiencing it is impossible from afar. Even perfect emulation can’t do it justice – there has simply never been a home controller that approached it. (I mean, the whole cabinet weighed 400 pounds!)
But if you are ever in an old-school-themed arcade, look out for Super Sprint (locations), or its two-player cousin Championship Sprint (locations) – and give it a turn or two.
Ivory is a Mastodon client, and their account switcher has a few interesting mechanics.
The standard one is that you can tap on your avatar, and get a menu in response:
But you can also drag down on the avatar, and get a different reaction:
This, I believe, is meant to be a slightly faster way. The settings option is gone to simplify, and the whole thing looks and feels more… gestural, in lack of a better word.
But I think it also serves one more purpose. The moment I saw this, I thought to myself “I wonder if I could just swipe on the icon itself?” and, lo and behold, this is actually possible:
Why does it matter?
I think for some power users of social media – perhaps people doing it professionally – you switch accounts all the time, and investing in this interaction being fast and smooth is important.
This whole small interaction system feels similar to switching apps on a Mac. You can choose an app in your dock with a mouse (the slow, but well-lit way). You can then learn to use ⌘⇥ and hold ⌘ to get to it quicker from a temporary menu. Eventually, you will also start tapping ⌘⇥ quickly, skipping any visible UI surface altogether.
There is also something great in seeing an interface that grows with you, or one where you can say “I wonder if…” based on your prior interactions and expectations, and the interface actually rewarding you for that thought.
Dragging and dropping files with your mouse in the Finder is nice, but there are many moments you wanna reach for the keyboard instead. Finder supports that. You dragged the file to a wrong place by accident? ⌘Z can undo it in a jiffy. You already have the right windows open? ⌘C and ⌘V can start a copy faster than the fastest of mouse gestures. And for cutting, there’s always…
…well, no, there isn’t. To the best of my knowledge, the Cut item in the menu never lights up around files or folders; you simply cannot ⌘X or cut in the Finder.
I think I understand the reasoning here. Of the Fantastic Four that is ⌘ZXCV, it’s ⌘X that is the truly dangerous one. Putting something in a clipboard is always a balancing act – a moment of inattention, and ⌘X turns into Delete. We generally accept it for text because the stakes are not very high, and text undo is basically too cheap to meter. Even then, occasionally, disaster strikes.
For files, the stakes are quite a bit higher. The creators of the File Explorer in Windows knew all that – yet, Cut still works there:
The way this is done is that the file isn’t immediately removed and put on the clipboard. No, only the intention of cut is registered – one half of the handshake, if you will – but the file stays put and nothing happens until the deal gets closed via a subsequent paste.
Or, almost nothing. There is a signal that ⌘X took effect – the cut files turn slightly transparent to indicate they put on their shoes and they’re ready to go. But if you don’t finish the paste, they become solid again the moment you cut something else, or when you restart the computer. (I wouldn’t be surprised if there’s a timeout, too, although I couldn’t verify that.)
I am not sure why macOS doesn’t do it the same way, particularly since they have a half-finished state designed already, for copying larger selections:
This feels like a strange omission I don’t fully understand, particularly given that it was vintage Mac OS that put the ⌘XCV combination on the map:
What’s even stranger is that the Finder has a command to move files – but an unusual one, buried inside the secret area in one of the menus, as an ⌥ alt to Paste:
It works. By putting the action on the other side, it avoids the “file floating in outer space” challenge, too. But it’s so undiscoverable that even though Move Item Here was introduced in 2011, I only learned about it a month ago as I started researching this more – and even though I know about it now, it still feels profoundly alien, being on the wrong side of the cut/paste continental divide.
In a reversal of typical state of things, it feels to me is that it was Mac that did a simple/cheap thing, and Windows that spent extra effort to dress it up into clothes of an existing Cut for reasons of familiarity, consistency, and better experience. Sure, a keyboard way to move files like Move Item Here is somewhat valuable, but nowhere near as valuable as a traditional cut, put under ⌘X ahead of ⌘V, working the same as everywhere else for years, even for decades, as one of the first denizens of motor memories all around the world.
I’m curious if you know of this pattern that existed for as long as I remember.
On a Mac, you can hold an ⌥ key (Option) whenever any menu is open, and often see more powerful, advanced, or faster variants of existing commands, helpful for a power user. Here’s Audio Hijack and Forklift:
And in the Finder’s File menu, both ⇧ and ⌥, and even the rare ⌃ (Ctrl) get to play:
This feels like a nice, thoughtful extension of the command system. It’s clever, too – the key to reveal secret things is the same key you would use for their shortcuts, so you can make the connection either cerebrally or in your fingers.
There are menus that treat it slightly differently, though. Here in Finder and Safari, you can see commands themselves changing when you press ⌥, no shortcut in sight:
(By the way, a nice touch in the entire system: Once the menu gets wider to accommodate longer strings, it doesn’t shrink on key release.)
And in other places, the modifier key only reveals an alternative shortcut for the same command:
I have mixed feelings about this feature.
On one hand, it has felt like dying art for a while now. Apple has never invested in the discoverability here, which I think kneecapped it. As far as I know, there is no hint these exist, and no way to see all of these options easily by clicking something on the screen. Even if you know the alternate name and you search for it, it doesn’t always reveal the key to press to see it natively:
It’s not as much fun to keep pressing all the modifier keys in every menu to find out what might be hiding in there. I wonder if an alternate version where any modifier key reveals all would be better? (And more compatible with commands without modifiers changing anyway.)
On the other hand, I really want this to spread its wings. In theory, the feature makes it possible to build up good motor memory habits, and understand some of the modifier key patterns: ⇧ often means “more,” and ⌥ (ironically) often means “alternate” (try ⌥ while selecting text on a Mac if you haven’t ever done it). It is also a clever way to accommodate power users without blowing up the menus for everyone else.
And, once in a while, I have a truly glorious moment – like what happened to me last year.
I scan a lot of documents, and merge them into PDFs using a database application called DevonThink. The process usually goes like this: I drag individual page files, select them, right click, and choose Merge:
But at the end of that process, I am left with the extra original files I no longer need, so I have to select these, and then delete them:
Not a big deal, but any “not a big deal” becomes a big deal if you have to do it dozens of times a month.
Yet, only after years of doing it this way, I had a thought. More precisely, my fingers had a thought. Surely, more people must be facing this issue? So, on a lark, I pressed ⌥ with the menu open, and saw this:
The precise option I needed was there all along, with a perfect label, hiding exactly in a place and under a key combination that was the first thing my fingers naturally tried.
I smiled so much. It’s a really great feeling when you’re rewarded for learning a pattern, when your motor memory does the job for you, and when you sense you’re on the same wavelength as software’s creators. Someone was there ahead of me and cleaned up this rare path for me.
A perfect example of breaking the “let me click” slash “the current action has the most momentum” principles we just talked about, in macOS Settings. When you try to add a new keyboard layout, the layouts you already have added are completely disabled, not just from addition, but even from selection:
On a surface this seems useful, right? A good signal to help you understand what you’ve already done? But this implementation also makes it unable to compare layouts of keyboards you have installed with the layouts you have not, since they previews are now in two different places. And this is important, as the differences between those layouts can be tiny, but significant.
This breaks momentum, and in the process also makes it hard to establish a good map of the entire space, since you are not allowed to walk around it freely. A tactical solution seems to be relatively standard at this point – disable the Add button instead, and indicate the already-added layouts in some other way.
Unrelated, but also – this popup-on-a-popup reminded me of this glorious old tram bezel city I’ve seen in Melbourne:
Ilya Birman on his blog talks about interfaces that unnecessarily slow people down, in a series of two posts.
In the first one, titled “Let me click,” Birman shows a few places that force the user to go through a roundabout series of clicks, instead of taking a direct route:
In Aegea’s comment settings, there is a “send by email” checkbox with an email-address field associated with it. If the checkbox is unchecked, there is no point filling in the field: the address is not needed for anything else. If the checkbox is checked while the address is blank, the system cannot send anything as it does not know the address. In short, these controls are interconnected.
Logically you could disable the input altogether if the checkbox is unchecked — there is no point filling it in anyway. But that is irritating. What if I want to enter the address and then turn on the checkbox? It would be even worse not to let me turn off the checkbox when the address is filled in. I want to turn it off — let me click!
In a follow-up, Birman talks about a specific example from the podcast app Overcast, whose creators faced with a tricky systemic challenge:
For any podcast, you can set how many episodes to keep downloaded on the device. Say you set the limit to three, then stop listening to a podcast regularly: the next three episodes download, and after that it stops downloading them — why waste the space? […]
[But,] what if someone manually asks to download an episode when they already have three downloaded? You can’t just immediately delete it to keep things tidy. […]
Finally, Marco’s wife Tiff suggests a solution to all the problems: just don’t let users download more episodes, she said, and show them a message, roughly: “Your episode limit is three, but this would be the fourth, denied”.
Marco liked the solution. I didn’t. Sure, it solves Marco’s problems, but not the user’s.
I haven’t seen the Overcast block in action, but all of Birman’s examples across both blog posts rang true to me. Motor memory wants what it wants, and stops for no one.
I wanted to add two things from my end.
1.
Birman presents this as “Let me click,” but I wanted to offer two alternative/complementary principles that helped me before:
Let me do things in any order. Here’s Google Home app, where I can change the temperature and the time for holding – but if I change the time first, it frustratingly resets when I subsequently change the temperature. It forces one specific order in an interface that suggests any order is okay:
The current action has the most momentum. The user is right there, active, tapping on things, wanting to get stuff done. If Overcast indeed throws a “your episode limit is three” message, then it forces the user to remember how to get to the settings and change it. The momentum is lost. The decisions of past me should not be as important as the decisions of present me.
2.
I agree with all the examples given by Birman, but things can be stranger, and sometimes letting people click or do things in any order can make the interface harder to understand.
In Figma, each text box dimensions could be set to be completely automatic (automatic width and height – used for short labels), with only specific width (and automatic height – used for paragraphs of text), or with manual width and height (used for graphic elements):
You can also see that in text boxes where width or height are automatic, the fields for those values are grayed out, and cannot be changed. In order to change them, you have to switch to a particular mode first, which makes the relevant fields active. This is to help people understand how those things relate to one another in a pretty tricky space, but it effectively means sometimes Figma won’t let you click, and will force you to do this in one specific order.
But this is only in those fields – you can always grab the object on the canvas and resize it, and it will switch to manual on its own, respecting the drag’s momentum:
The inconsistency here is intentional. I’m not saying these are the right choices, and we indeed heard from some users frustrated that they cannot click easily when they already understand the system (unfortunately, I am aware of no good affordance for “breaking” a disabled field in modern GUIs, like a molly guard). I mostly wanted to share that these things can be hard, and the balance between “let me click” and “not being able to click can be helpful in understanding the system” tricky to achieve.
I keep thinking of the story from a decade ago when someone’s phone rang at the front row of a New York Philharmonic’s concert, prompting an actual performance halt, and anger from both the conductor and the audience. The ashamed patron’s eventual explanation was: “I turned the phone to silent, but I also had an alarm set up.”
Let’s assume this is actually true (in people’s reports, the phone rang the marimba ringtone, which isn’t standard for alarm – but people’s recollections are routinely flaky, too). This, I think, is a perfect example of “let me click” in action. You can set an alarm for a certain time, and you can subsequently turn the phone to silent. You can also do it in the opposite order. You just gave your phone two inconsistent instructions, and the phone logic decided in either case the alarm will win.
I can’t think of an easy interface solution here. It feels correct for the order to not matter here. It’s hard to imagine disallowing you switching to silent mode with any alarms on, since the silent mode also affects calls. It’s also hard for me to imagine any effective UI warnings at any given moment in the process. (Besides, back in the day, iPhone used to give you a gentle one via a little alarm icon visible in the top bar.)
The silent mode and the alarms are, as Birman put it, interconnected – but their connection is ultimately tricky to explain to the user.
I believe every modern operating system allows you to set the key repeat rate and the delay before first repeat:
By default, on a Mac, the initial delay is 500ms, and the key repeat rate 1000ms. You can adjust both:
Initial delay goes from 250ms, to as long as 2s.
Repeat rate goes from 2000ms (once every two seconds) all the way to a whopping 33ms (30 times a second).
On a PC, the shortest values are exactly the same, but Windows 11 only allows up to 405ms between key presses, and 1000ms of initial delay.
My suggestion is to make them both as short as possible for testing.
Why would that be helpful? There are two reasons.
First, it’s good to test whether your keystrokes behave well when repeated. Occasionally you might want your interface to suppress repeating, or do something special if the key is held longer.
Second, it’s a good test not just of repeating, but also of a user pressing keys really fast. It’s very important for the UI to never make you wait for any animations or transition, and a lightning fast 30fps repeat rate is a good way to stress test your system this way.
Here’s me holding the Tab key in Figma at a standard key repeat speed:
It’s nice to see those transitions help you orient yourself as you’re thrown around the canvas. But look at what happens when I do the exact same thing with the key repeat cranked up:
The transitions are now slowing the UI so much that you can no longer see any canvas movement – the canvas only catches up when I release the Tab key.
The smooth movement assumed a certain minimum key rate, and perhaps wasn’t tested with a faster one.
There are of course ways about it, like speeding up the interactions, suppressing the transition smartly, or introducing a custom repeat rate if needed – I talked about it a bit in my essay about designing fast keyboard interfaces.
Here’s an example of an interface that suspends transitions at a fast keyboard rate and responds in real time:
This looks very chaotic especially since you’re not in a driver seat, but this interface is at least honest and doesn’t make the user wait. And it might not be pure madness – you would be surprised how fast people’s brains and fingers can react.
But first, you have to be aware of the problem, and I think setting up those key repeat values as short/fast as possible will help you find more of those kinds of issues.
(And you might choose to keep those high speeds anyway in regular use, like I do.)
Oh damn, you caught me in the middle of something. I was just trying to make a list of Windows versions for a friend – in Google Docs, of all things.
Should be easy. I already grabbed this one off of Wikipedia and massaged it a bit, but it’s still kinda ugly:
I don’t love the tight padding here. Let me try something bigger, like 0.08 inches? I’ll just select the first column and punch the number in, and…
What the hell!!!
Jesus. What happened? Okay, let’s press ⌘Z to get out of it…
Oh, no. Maybe I can press Esc…
Shit.
2.
Okay, here’s what happened in precise detail:
I pressed Backspace to delete an existing zero:
I started to type “.08”. Pressing “.” was okay, although it showed a pretty thirsty tooltip:
upon adding “0” Docs removed the “.” and so the output was just “0”:
upon adding “8” Docs removed the leading zero and we ended up at just “8”:
the page saw the resulting “8”, applied 8 inches of padding (a hundred times of what I wanted!), and just showed it to me in real time, resulting in a profoundly unrecognizable table that looked like something went horribly wrong,
pressing ⌘Z didn’t do anything,
pressing Esc only removed the focus from the input field.
I believe I can explain exactly the chain of reasoning and bugs here:
To start with, any number entered is immediately previewed on the left. This generally feels good!
Instead of fixing my input on commit (Enter), the input is being rewritten on the fly, as I’m typing. This wouldn’t be my recommendation, but I can understand this design decision.
If I typed just “08,” it would be rewritten to “8”. This makes some sense since it normalizes the numbers, and makes them consistent.
If I typed “.1”, it would be rewritten to “0.1”. Sure, fine, a similar idea.
However, typing “.0” rewrites it to just “0”. I believe this is a bug or a lack of imagination – it should be rewritten to “0.0”.
⌘Z doesn’t work. I believe this is a bug where system’s rewrites of numbers don’t put the change on the undo stack. (As a matter of fact, it appears worse than that – pressing ⌘Z a few more times ended up rewriting my number to “80”, which would have made this even worse!)
So, in effect, my “.08” got mangled to “8”, applied immediately, and wasn’t easily undoable.
If just one of these bullet points below behaved differently, I would not end up in this situation:
with numbers not being rewritten on the fly, there would be no problem,
with “.” not be aggressively rewritten to “0.”, there would be no problem,
without live preview, the rewrite could’ve been caught and fixed it by hand, instead of panicking seeing a huge change on the screen that felt like data loss,
with a fully functioning input field undo, the moment of panic could be reverted,
with Esc to abort instead of commit, likewise.
All of these decisions made sense and didn’t feel dangerous or important in isolation. Together, the holes in cheese aligned perfectly, creating a pretty scary experience.
I spotted this kind of a keyboard shortcut pattern the other day. Here it is in Photoshop:
Here it is in DevonThink:
And here in Linear:
Those “ordinal” keyboard shortcuts feel nice and orderly, and look so elegant, too. But beware! The moment you’ll want to introduce a new option, or even reorder the ones you have, you’ll be in trouble: Either you preserve people’s motor memories, and then the elegant ordering immediately goes to hell – or you will have to change an existing shortcut to something new, and frustrate your users. Might be best to be really confident in your selection being forever locked before attempting this.
But then there’s a similar treatment here in Linear:
Or here in Raycast (when you hold ⌘):
Or here in Ghostty:
Those three look like the same idea, but they worry me less.
Why? Because these shortcuts more clearly point to a position rather than a thing. Here, the mechanics of the UI themselves convey that people, commands, or tabs are going to be moving around. Of course, some will get used to “1 means assigning to Marcin,” or “⌘3 means the Unsung tab” – the same way we get used to “second item on the Recent list” or “at the top of the third search result page,” if they start repeating as a pattern – but at least it feels to me that the interface here is more honest about what it can promise.
I often think about this, by the way: Not what the interface conveys in the moment, but what it promises in the long run.
The most common use of double-clicking is as a shortcut way to perform an action. For example, clicking twice on an icon is a faster way to open it than clicking once to select it, then choosing Open from the File menu; clicking twice on a word to select it is faster than dragging through it.
I knew that double click an icon was a shortcut to the first action (typically Open), but I never really thought of double-clicking a word as a faster way to drag across to select it – even though, in hindsight, it makes perfect sense.
Another vintage thing I learned of recently from a coworker is this, also covered in the 1987 HIG:
If the user begins a double-click sequence, but then drags the mouse between the mouse- down and the mouse-up of the second click, the selection becomes a range of words rather than a single word.
This doesn’t feel (to me) like a very pleasant gesture to perform repeatedly, but what feels nice about it is that it automatically snaps the selection to the endings of the words:
Part of me would prefer this to be the default behaviour when selecting more than 3 words, or so, so you could be less precise.
Anyway. The double clicking to perform default action applies to a lot of lists of things. Here are some examples from Scrivener, Word, and Lightroom – you can double click on each of these items to proceed, without having to select and click the button:
But sometimes the creators of such dialogs forget. Here’s Screen Sharing in MacOS, and a notification in Chrome where only the slow path is available – double clicking on items doesn’t do anything:
The tricky part about not being a good citizen of a shared user interface is that those omissions aren’t just local to your app – they can ruin the gesture in other places, as people’s fingers learn to distrust it in not just your app, but in general.
A strange thing happens when you press Enter on an empty item in a list in most text editors – the entire list item disappears:
This feels counterintuitive. Isn’t Enter for committing and adding more things? Wouldn’t Backspace be the right key to press to break a list?
Yes, and no. I am not sure who invented this pattern (I spotted it first in Word 95), but that someone understood a strange interaction contract existing in text editing – Enter is actually an escape hatch. In text editing, no matter where you are, you can always press Enter multiple times to just create more room for writing.
In an app that doesn’t cancel a list on Enter, you can face a terrifying moment where you get stuck in a list, and getting stuck is never fun.
This principle feels so useful that I see more and more apps apply a version of it for other things. For example, in many modern text editors pressing Enter after a headline returns you to regular text, just so it’s not as easy to get stuck in a headline style:
Nice moment in Slack and Medium – when logging in, the login code that arrives via email is already there in the subject, in addition to hiding inside:
This feels good for two reasons. One is the obvious one: you see the code earlier, it might show up in a notification, etc.
But also, this should prevent multiple login codes to be threaded as a “conversation” inside your email client, since threading is based on the email subject – and threaded utility emails can be extra confusing.
A computer science professor Paul Cantrell, on Mastodon:
Creative work keeps taking roughly the same amount of human labor / attention / care, even as new technologies accelerate or remove things that used to take time.
This is because creativity is fundamentally not an efficiency problem; process is not just the means of producing output, but rather a labor vessel that holds the near-invisible work that is truly important.
One can feel the care that goes into creative work without being aware of that work, or even being aware that work of that type exists at all. This feeling is approximate, loose, vague, but cumulative and eventually all-important; work with no care behind it wears thin and tends to fade as people live with it over time.
I constantly see some people praise it not for what actually makes it good, but by taking the things it’s bad at and turning them into a puzzle to have “fun” solving.
I’ve had people tell me how “fun” it was to build a macro to handle some one-off text-refactoring problem. But when I looked at what they were doing and how long it took, my honest reaction was: I could have done that in Sublime in a minute with multiple cursors, or just written a quick script. […]
That’s what I mean by “invisible tools”. When you’re proficient with your editor of choice—whatever it is—it disappears into the background. But the moment it cannot handle something easily, it stops being invisible. What baffles me is that so many people treat that friction—the effort of working around a tool’s limitations—as the “fun” part, and then advertise it as evidence that the tool is great. […]
The text-editor-macro anecdote I mentioned is really about a gap between feeling productive versus being productive. There’s a sensation of cleverness that comes from solving a fiddly problem, and it’s easy to mistake that feeling for actual output. A tool that makes hard things feel heroic and clever feel like an achievement can register as “powerful” while quietly being slow. The honest test isn’t how engaged or clever you felt, it’s wall-clock time and how many mistakes you made getting there.
This I had more of a mixed reaction to.
I think it’s necessary to expect from tools to get out of the way, but there’s also nothing wrong with having fun with them.
My simple go-to example is this: When writing code, I sometimes use Find & Replace All, and am done within a few keystrokes. But sometimes, I press Find and then replace one at a time, jumping methodically through the file, and seeing each string in situ before changing it. I know the tool could do it all for me. I know I could be more efficient. But this intentional slowing down allows me to refamiliarize myself with the code, visit its forgotten nooks and crannies, and make sure I understand where and how the thing I’m changing is actually used.
The editor I use allows me to not be efficient when I choose not to be. In my work, flow operates at different speeds; a good tool understands that and doesn’t force me into a particular one.
I think ultimately indeed, the tool does need to disappear, and make you be in charge of whatever speed you want to operate at, and how much friction or difficulty you choose to face (do you bump the lamp or not?). But it’s not as simple as always “reducing wall-clock time and mistakes.” Like Cantrell says above: Creativity is fundamentally not an efficiency problem.
A typical use of Safari means two groups of sites: a regular set on the right, and “private” pages on the left (this is what Chrome calls “incognito mode”):
As expected, you can tap on either label, and switch to the relevant group with ease:
It also feels like you could slide it – and you can, except…
…you immediately encounter a Scroll Lock problem. You are not dragging the pill – you are dragging what’s underneath the pill. To switch, you have to go the other way:
You can immediately intuit some inherent unpleasant complexity of the whole system – not just in it “going the wrong way,” but also in how it creates room and then contracts it, in two separate steps, after you’re done.
The reason is that you can actually have more site groups than just the initial two. You can even drag to where the new site group would be, and create it this way:
I normally welcome these kinds of accelerators. But here, this feels overdesigned and confusing, as if someone drugged the tab instead of dragging it. The very same natural gesture – a left swipe – that should feel safe and send you to Private, will now put you in a scary new full-screen/keyboard-out flow you almost never need.
Why this relates to system design is that on/off toggles in iOS were recently redesigned to resemble oblong pills:
Those do respond to dragging as you’d expect:
Along the same lines, on the springboard pagination pill, dragging to the right means the next page:
And so now the system is schizophrenic and identical-looking design primitives mean the opposite things. It’s as if the computer itself kept randomly pressing Scroll Lock for you, preventing you from developing a solid understanding of the system first, and motor memory second.
I think the mistakes made here were twofold. First, the design overoptimized for two unnecessary things: people actually using site groups (not common), and ease of use in creating site groups (not important). My slightly cynical hypothesis is that this design presented really well in demos, which sometimes can derail a project. A more cynical theory is that this led to “accidental discoverability” that made the site group metrics look better.
Second, and more important part: This particular design received an exception that it didn’t deserve. No one noticed the systemic challenge of similar UI elements doing opposite things, or people who did were not effective in pushing back. The metrics for new feature discovery are easy; the metrics for user confusion or frustration do not usually exist.
This is how interaction systems slowly fall apart. As I mentioned in the first part, it is likely that Safari’s exception will now be treated as “blessed,” and start spreading further. Given enough time, more and more pills will go in whatever direction they want when dragged – and people will learn not to trust any of them.
I know the feature is actually called “tab groups” but I called it “site groups” intentionally, because otherwise it’s a tabbed control controlling tab groups, and things get confusing really quickly. Also, thank you to Martin Hoffman for initiating this post.
As a teenager I adored Norton Commander, learned some UI magic from early videogames, and dreamed of a Mac my family couldn’t afford. But I think the first product that taught me something truly memorable in terms of user experience was, of all things, Excel 97.
Excel 97 cooked. It was rendered in brand-new, gorgeous Tahoma, had keyboard underlines everywhere, and felt complete in terms of features. Sure, you could maybe sense the beginning of the bloat with the reorderable menus and the bolted-on Clippy, but Excel 97 felt tight and zippy even on an abominable computer put together on a high-schooler budget from used hardware pieces that were never meant to meet.
But what made me love Excel 97 was one specific thing: the Repeat command. You could invoke Repeat via Ctrl+Y, but my fingers learned to love its stranger alter ego, F4. This is what pressing F4 did:
It was a bit more clever than just replaying the previous action. If you changed the font and its size, it would repeat both:
And it didn’t only repeat formatting, but also some actions, like inserting new rows, or merging cells:
There are, of course, many other solutions to make these things less tedious:
make a complex selection first (if possible), and format later,
use styles instead of naked formatting,
copy/paste formatting only,
paint format from one cell to another,
record and repeat keystrokes,
save individual actions into bigger macros and reuse them.
…and some of them were even available in Excel 97.
But it was Repeat that stole my heart. I think this is because it was two types of magic combined: the magic of motor memory + the magic of smart software.
The first part meant that soon, it didn’t really feel I was pressing F4. The shortcut lodged itself in my fingers and in time, it became a gesture as natural as pressing Backspace or arrow keys. All the other alternatives above required thought or bureaucracy, but Repeat was mindless – I would occasionally watch my hand perform it on its own.
The second part is that F4 felt like it was reading my mind. There were no options; Repeat just always seemed to do the thing I expected from it. As I understood this kind of functionality more, years later, I learned to appreciate the deeper thinking necessary to make it work. Repeat wasn’t actually repeating keystrokes and commands – it was some tricky dance of adjectives pretending to be verbs, with groundwork necessary so that the system felt stable and consistent.
The feature wasn’t flashy. It didn’t even have a separate toolbar button. But it was beautiful.
Repeat is still there in Excel on my Mac today, in Google Sheets, and a few other apps like BBEdit. It’s often an extension of redo, although I prefer thinking of it as something independent. Some apps like Sublime Text or Keyboard Maestro have a quick record/repeat function that feels similar, although it’s not as smart – this one is solely repeating keystrokes, and requires you to start recording first. (Part of Repeat’s beauty is that it works retroactively.)
If you look at the keyboards on my desk, they all appear blank…
…with one exception:
Repeat was fantastic. I wanted it as a user elsewhere. Then, as a designer, I wanted it for my users.
In time, I learned to appreciate other things like it, but both the solitary legend on my keyboard, and the current favicon of Unsung, are an homage to the original, 1997’s Repeat. (Original to me, at least. Repeat existed in Office 95 too, and perhaps in some other apps before that.)
If I got to spend my entire professional life designing hard-to-make, invisible things that make other people feel powerful and awesome, it would be a life well spent.
The calculator app has been preinstalled on iPhones ever since their debut in 2007. For the longest time it hasn’t been anything more than a standard four-function calculator with a decades-old feature set. If you’re not careful, however, you can mess up even that.
Ten years into iPhone’s history, iOS 11 introduced a problem just like the Nothing Phone rotation – quickly tapping on keys would show them as responding, but the actual action wouldn’t be registered. Michael Tsai’s aggregator’s first entry has a video from Stephen Heaps:
It shows typing 1+2+3+4 where iOS forgets one press of +, resulting in 1+23+4 = 28. Many more people posted about it afterwards, and showed various other examples.
It is oddly enthralling to see a computer fail at basic math. But what’s particularly historically interesting and perhaps even more embarrassing for Apple is the absolutely rich history of solving this kind of a problem.
Calculators evolved alongside typewriters as the earliest devices with button-like (as opposed to piano-like) keyboards. But the stakes were different.
Imagine a badly constructed typewriter and all the ways it can disappoint you: the letter might be faint if you press the key lightly or puncture the paper if you press it too hard, the output might be misaligned, or the typebars will jam in some way, forcing you to go again.
A typewriter has to work hard to divvy up a blank, analog piece of paper into a reliable, pleasant-looking grid via escapements, ratchets, and so on. But a calculator’s work to convince the analog world to be digital was much more important. After all, it’s not likely that the typewriter key you pressed will output the wrong letter – but on a badly constructed calculator, a light press of 5 could absolutely output 4, or 6, or 4.5.
And, while the typewriters only take your words verbatim, the calculator’s job is precisely to create new numbers out of the numbers you type. An imprecise mechanism can mess up that math. A jam could perform a partial or nondeterministic calculation. Adding 1 to 999,999 and the force necessary for the resulting cascading carry could break a device in the middle of work.
On top of all that, languages have a built-in redundancy. Evn if yuo mak many typoes, th sentece can stil be understod. But all numbers basically look alike. A calculator could make a mistake when it comes to a number that is absolutely vital for your payroll, for engineering, or for navigation – and you would never spot it.
Understanding all this, many calculator makers even already in the 19th century spent a wild amount of effort convincing people not just that their devices were helpful, and fast, and easy to use, but also that they could be trusted. Buttons were carefully weighted. Comptometers came with a locking mechanism. If a machine felt something didn’t go right, it would stop working and require a hard reset. The message was: “You can trust me, because I won’t ever show you bad math, and I’ll stop myself before I will ever lie to you.” Charles Babbage was so confident in his Difference Engine that he welcomed people to try to mess with its mechanical wheels in the middle of the calculation, convinced that even a sabotaged machine won’t ever make a mistake.
Just like with the Selectric decades later, those things were solved by people who cared, in the much harsher mechanical conditions.
Of course, I don’t expect everybody at Apple core iOS team to be a calculator UI historian (although it would be nice for at least one person on the team to be one!). It is embarrassing that no one on the team had enough imagination to realize that making a button respond to a quick press during animation, but not register it would cause all sorts of serious trouble. (The bug was fixed in iOS 11.2 by removing the animations, and subsequently the animations were brought back without the original problem in iOS 11.3.)
But maybe the bigger embarrassment is that Apple didn’t have a battery of tests to run on top of the UI at various speeds, mimicking fingers of what must be millions of people using the calculator app. That, too, has been a standard procedure for decades.
Those tests seemed missing in 2017. I hope 2+3+4 years later that’s no longer the case.
Perhaps ironically given the subject matter, I found this 34-minute video by Razbuten a bit intense, but I would still recommend it to people who work on onboarding, settings, etc.:
In the video, the author tries to answer the question: how to make any given game a challenge, given there is no universal standard of difficulty and every player arrives at a game not just with different skillset, but also likely different goals.
There are many techniques a game can use to adapt to the player – a simple upfront difficulty selector, complex difficulty settings, a training level, adaptive difficulty, accessibility/assist modes – but there are no easy answers. Each method comes with pros and cons, and perhaps the very notion that a game should adapt to the user is flawed; some players might find it more rewarding to have to step up to the game instead.
In the video, Razbuten covers a lot of examples really well. I’m not going to say any of this maps 1:1 to productivity software as goals of games are very different than goals of apps… but even though I have never played any of the games mentioned, the examples made me think. After all, some of the psychology of mastery will be the same between these two realms. (I bet there were at least some of you who saw the previous post about LaTeX and thought “this looks hard and fascinating – I’m going in,” and others took a note to never approach it.)
Collier talks about why physicists prefer LaTeX to Word. LaTeX is sort of a nerdy HTML that predates HTML. It looks like this…
…and given how nerdy HTML already is, you might imagine this is a power-user tool that’s chiefly about power and control. But Collier makes the argument that there are some things that LaTeX makes much easier:
there is absolutely no need (or peer pressure) to spend time styling the document by choosing fonts, colors, etc.,
there is no “live preview,” and making a PDF is a separate step similar to compilation in coding – which means it doesn’t constantly occupy your mind,
GUIs can slow you down because the keyboard is faster than the mouse,
(Of course, there is also the issue of typographical craft of LaTeX documents set in Computer Modern, but let’s save this for another time.)
Also, the video starts with Collier apologizing for potentially making the audience feel dumb in a prior video. I don’t think it’s a joke, and I found it thoughtful and refreshing.
One thing I was (and still am) worried about when it comes to my recent big interactive essay is that by showing all these classic desktop examples, the whole thing might appear old-fashioned, relevant only to a bygone era.
Yet, the challenges it shows are universal. Here’s something I just spotted. This is how you rotate an image on an iPhone and on a Nothing Phone:
It’s a pretty standard control – tap once to rotate counterclockwise, tap a second time to do it again, etc. – with a helpful transition of the photo’s orientation so that you don’t lose yours.
Now, I’m going to exaggerate the problem a bit and tap 90-degree rotation quickly eight times. Eight times should result in what engineers call a “no op” – the image rotating twice in full, and ending up where it started. That indeed happens on the iPhone:
But it’s a different story on the Nothing Phone/Android:
iPhone will remember and buffer the taps, so that the second, pending rotation will happen as soon as the first is done. The Nothing Phone button gives you a tap confirmation via both haptics and sound, and then ignores the tap if a previous rotation is still animating.
Why does it matter?
I often keep thinking about the framework of situational disability, stating that disability is not just something that happens to a few people and no one else. No, pretty much everyone will occasionally encounter a situation that will make them effectively disabled, and this is why accessibility matters much more than many of us assume:
I think similarly about casual and non-casual use. Photo-taking on phones is typically casual. Phone cameras are typically very good at detecting the photo orientation – but get confused when you’re pointing down. Now, as an example, if you had to take photos of a bunch of landscape documents, you might end up having to rotate dozens of photos, one by one. And it would be so much more predictable and pleasant if you could just tap the button three times at any pace you wanted without thinking, without paying attention, without getting your UI blocked by an animation that no longer helps you.
This is, I suppose, “situational power user-ness.” Given a long enough timeframe – or, in this case, a large enough population – even a casual interface like phone photo editing (or, GarageBand) will meet someone who will have no choice but to treat it more seriously and expect more from it.
By the way, buffering the taps is not the only answer. You can also stop/accelerate the animation after an interrupting tap, and it seems the iPhone does that as well. But the rule is: never force the user to wait for the animation to finish.
Here’s another nice detail. If you press and hold ⌘⇥, you will eventually stop at the end. (You can then press ⌘⇧⇥ or ⌘` to go in the other direction.)
However, if you are already at the end, pressing ⌘⇥ again wraps around to the beginning:
The issue of whether to wrap around or not is more universal; you can see it in many lists, ⌘F, and so on. On one hand, it’s nice to have a solid deterministic stopping end that you can rely on, especially since sometimes the last item on the list is special (“See more items…”). On the other hand, going all the way back from the end can be frustrating, too, especially on a Mac that does really strange things with Home/End/PgUp/PgDn keys.
I thought the hybrid approach that ⌘⇥ is doing here was clever, and might be applicable elsewhere.
This is what happens when you go to the homepage of Gemini and start typing quickly:
Mechanically, I think this is React or some other framework setting focus again with some delay, but the end result is… rather disturbing.
While the technical solution would be to fix the problem or at least do not set focus again if already set, I wonder what’s the real challenge here. I imagine it might be that the testing process (if any) assumes using the mouse or trackpad first. In this case, moving the hand to the keyboard to start typing gives the interaction just enough delay to miss the second, unnecessary focus.
I think a good assumption to have for all common interactions is that for some users, fingers are already on the keyboard and things can happen so much faster than you expect.
Not accounting for that, the creators of this flow inadvertently broke one of the cardinal rules. We talked about it in the context of mouse pointers before, but it applies as well to text: don’t move my cursor for me.
There is nothing quite so frustrating as a persistent user interface papercut. You know it’s there, but you keep running into it because the moment you start thinking about what you’re doing instead of how you’re doing it, that knowledge slides away until BAM you run into it again.
I think this is really nicely put and highlights about why it’s very important to care about this kind of stuff.
If you forgo a standard interaction out of carelessness, a bug, bad systems thinking, or for other reasons, you’re not just making your users frustrated by something not working. You’re also at risk of making them frustrated atthemselves, assuming they can change what their fingers do easily, not fully knowing that a) this is motor memory, not just regular conscious actions (and any memory is hard to “update” intentionally), and b) motor memory is separated from regular, declarative memory, and not possible to reason with using the same techniques.
(As an example, it’s very hard when keyboard shortcuts or mouse gestures disagree between apps, because while you consciously might know which app you’re in, that’s not necessarily true of your fingers.)
Waider continues with an example:
The canonical example of this, for me, is Microsoft apps on macOS: even now, decades after Microsoft started producing macOS versions of their apps, they insist on largely disregarding the native UI idioms in favour of their own. Current pet hate is that if I’m commenting on a document, the Ctrl-A/Ctrl-E actions do not work, and boy howdy do I use those constantly.
My recent example is that even though I wrote about Safari overriding the natural “scroll to top/bottom” tap gesture on their tabs – so I am aware of it in my declarative memory, I know Safari designers messed it up, and I know exactly what to do and not do – my fingers still occasionally tap to scroll in Safari anyway.
Every once in a while, I stumble upon a long thread in a random corner of the internet where someone discovers Paste And Match Style, and everyone erupts in applause. “Yeah, it’s a life saver.” “I use it all the time.”“I can’t believe this isn’t the default!”
Then, inevitably someone chimes in: “Oh yeah? I can show you how to make it the default.” And they explain how to wire ⌘V to use Paste And Match Style.
And I always get worried seeing that.
I believe this is the core problem people are bothered by before discovering PAMS – when you copy and paste from another doc, you inherit its style/visual appearance:
And Paste And Match Style, well, does what it promises:
This feels nice. So, what’s the problem? The problem is that PAMS is drunk with power and flattens everything on its way:
That includes:
emphasis by italics or bolding
links
bulleted and numbered lists
strike-through text
headlines
None of these are “style.” This is actual information that should not be removed. If you wire PAMS as your main ⌘V shortcut, or even if you use it occasionally, you might remove valuable data from text you’re moving around, without even noticing.
(And if you do notice, the frustrating irony is that recreating the information lost in transit – for example, re-linking things one by one – is often more work than fixing the style would be.)
If you are designing an app that handles rich text, here’s what I have seen others do:
Do not have styles to begin with. If you use Notion, Dropbox Paper, Medium, or anything that relies on Markdown, they give you no way to customize fonts, colors, letter spacing, and so on, so regular superliteral Paste has a limited blast radius and works well:
Have a very strong center of gravity toward the default style. Apple Notes does this well. Use Notes for years, paste into it from all over the world, and you might never realize it allows you to change fonts and colors. Its default Paste removes style, but it doesn’t remove any valuable information like links or bullet points.
Notes also introduces a shortcutless Paste And Retain Style as a third option after a “semantic” paste (which keeps data and removes style) and PAMS (which removes everything), for those who really want to paste extremely literally:
Word has Paste And Match Formatting that seems to be what Notes does by default, but it’s not the default:
Help users understand the options they have more. For example, Word offers a little post-paste menu. I don’t personally love (it doesn’t have a preview + it doesn’t remember my preference + the options are scary), but it uses better-than-default language like Keep Text Only, and it protects people from the harrowing backrooms of its own Paste Special:
Have some contextual rules – for example Figma does things differently depending on whether you paste into a new text box (preserve style), or a text box that’s already filled (match formatting).
(If you’re seeing some other apps doing something interesting, please let me know!)
Doing the right thing won’t be easy. Books have been written about the illusion of the difference between “stylistic” and “semantic.” People use bolding for either. Others treat headlines as visual style, right aligning means something different in English than it does in Arabic, you might still have to normalize indentations, and so on.
But I believe it’s necessary to put in the effort to make regular Paste work as well as humanly possible, rather than relying on people to know about the far-from-perfect ticking time bomb that is PAMS.
The theme: What does it take to build interfaces that truly allow for fast operation – and why that matters.
If you like the interactive details posts here on Unsung, the essay is kind of a concentrated dose of all that. You can technically read it on the phone, but it’s so much better on a computer (or a big tablet). It has ~40 interactive playgrounds, and sounds, and a glossary, and all sorts of fun stuff I’m doing for the first time.
I wanted to share some things I learned over the years, and nod toward mostly anonymous creators of UX inventions I’ve long admired. I also thought it could be interesting to make interfaces appear as machinery – you’ll see what I mean.
You might have seen Bret Victor’s 55-minute Inventing On Principle talk soon after he gave it in 2012. If not, you should check it out. If you did, you should check it out again and see how it makes you feel today:
It is about interactions but in the service of something grander, which (if I’m doing my job well) you might recognize as Unsung’s core theme.
Victor – a designer, researcher, and computing historian – gave a few other talks in the few years since, and I thought a little guide might be helpful:
Drawing Dynamic Visualizations (35 mins) is specifically about information visualization, chiefly a demo of the “Illustrator, but programmatic” tool showed briefly in the above talk. There’s also a bit more theory.
Similarly, Stop Drawing Dead Fish (53 mins) is a demonstration of a different programmatic tool to make animations.
I love this blend of theory and practice, inspiration and pragmatism, high- and low-level. The tools look surprisingly professional for research projects, but underlying their microinteractions is a deep philosophical stance. It all reminds me a bit of Jef Raskin and Doug Engelbart.
Victor’s last talk of this era is Seeing Spaces (15 mins) from 2015, serving as a sort of introduction of him moving toward computing in physical spaces. As far as I understand, Victor has been spending time on Dynamicland since, which is definitely more physical computing, but also a lot more academic and scrappy, and as such out of range for this blog.
(His website is worth checking out, especially if you’re not in the mood for talks and would like to get to know his work in a different way.)
It’s an impressive list that garnered universally positive reactions, but I have one observation:
“Fast” and variants thereof appear on this list 59 times.
“Reliable” and similar words appear 22 times.
It’s true that everything could be faster and a whole many things should. Speed is paramount to great user experience. Speed is also more than just speed; there are nonlinear aspects when latency or delays cross invisible thresholds that can drastically change app usage for the better.
But in my experience, much more often the things that frustrate me about using Apple’s products are not issues of speed, but issues of reliability:
I don’t need faster network connectivity in Finder, but I struggle with computers not appearing, a pizza cursor that occasionally just dies spinning, or randomly being thrown to the root of the networking volume.
I don’t need AirDrop to be faster. I just want it to connect reliably every single time, give me consistent and understandable UI feedback, and stop forgetting I’m not just “everyone” when sending stuff to myself.
I don’t need Messages to be faster. I need them to just, you know, not haphazardly stop syncing across computers on occasion.
It’s not just me. 15 out of 17 bugs listed on the Bugs Apple Loves site are about reliability. None seem to be directly about speed, although more on that in a second. Or, here’s a recent list from Ilya Birman – different issues, but a similar shape.
I’m going to say it: Speed is an easier problem. Not easier in an engineering sense; I’ve seen an engineer try to carve ten milliseconds out a busy computer’s schedule, and in that moment, one must truly imagine Sisyphus happy. But it’s easier as a problem: it often comes with a lot of pre-built telemetry, plus a clear goal of “here’s Xms and X now needs to be smaller.”
Reliability is much harder, more difficult to debug, reproduce, agree on metrics for, even find ownership of – just generally fuzzy around the edges, and less obviously thrilling as a challenge. It needs more champions and structures.
I know a simple marketing slide is not meant to be an accurate representation of Apple’s efforts. I wasn’t at WWDC so perhaps the vibe in the room was different than what this slide represents. Yet, the slide exists and I’m allowed to judge it.
(And yeah, I know in some cases speed and reliability are correlated. After all, if you have a timeout, making something finish faster and do so before the timeout will turn it from unreliable to reliable. But hey, I wasn’t the one choosing the words on the slide.)
I just… I would be a lot more excited if the 3:1 ratio of fast-to-reliable on that slide went the other way.
In game development, there is this strange effect known as “tunneling.” It happens when you do collision detection. Imagine a simple situation where every time you move a cube, you also test whether it touches the wall – and if it does, you make it bounce off of it.
This works great, but if you move and detect the collision less frequently, something weird can happen:
Here, the movement was so coarse that there was no point at which the cube touched the wall, so the collision wasn’t detected, and the cube passed cleanly through… as if it made a tunnel.
The easy answer seems to be “well, run the collision detection more often then,” but… how often? And what if the entire game engine runs off of computer’s frame rate, which you are not in control of it at all? All in all, it’s not a trivial challenge, although various techniques exist to remedy it.
We talked before about another challenge with frame rate dependence. They’re not limited to games; for interfaces that are based on physics, they will rear their ugly head, too. But tunnelling happens in simpler UI situations as well. Here’s an example from Photoshop – I’m holding a button and if I drag slowly, each item will be toggled. But when I start moving fast…
Fortunately, the remedy here is much easier than in the complex world of videogame physics: just remember the last one touched, and toggle everything in between that and the current one.
Pointer input isn’t continuous. The browser reports the pointer’s position as a series of discrete samples, and when you move fast, those samples can land far apart. Flick your wrist and two consecutive pointer events might be a hundred pixels from each other, with three small shapes sitting untouched in the gap between them.
A naive eraser asks “what’s under the pointer right now?” on every pointer event. At slow speeds that works fine. At high speeds it tunnels straight through anything that happens to fall between two samples, which is exactly when people use the eraser most aggressively: big, fast, careless swipes to clear a region.
Here’s Ruiz’s example of the eraser in FigJam, with heavy tunneling:
And here’s one in his drawing tool:
You can click through to learn more and see the algorithms, but either way it’s worth remembering: if it applies to your interactions, think about the “flyover states” and make sure to make things deterministic, regardless of whether the mouse is a tortoise, or a hare.
And, I liked Ruiz’s sentiment at the end:
[…] The decision to test segments instead of points is the difference between an eraser you can trust at speed and one that mysteriously leaves survivors behind. Users will never notice it working, and that’s the point.
In 2021 and 2022, product manager Steven Sinofsky wrote a…
…first-person account of what I saw at the PC revolution from the perspective of joining Microsoft as a newly hired software design engineer fresh from graduate school working on developer tools, through my time as a program manager and ultimately leading Office, and then moving to Windows, and everything in between.
The first part covers the challenge of the team in 2007, taking stock of Office after almost 25 years of its evolution. (Number of toolbars in 1983: one. Number of toolbars in 2003: 31.) The second part shows great screenshots of all the Office versions from 1.0 until then, and the remaining four cover the Ribbon redesign process.
Regardless of how you feel about Microsoft Office today, and whether you consider the Ribbon interface a success, it’s a perfect weekend read as it covers universal challenges of software complexity and change management.
It’s such a potent series I’m sure we’ll come back to it. It covers a lot, including – in the first part – wrestling with a definition of bloat or complexity, which in the context of Office was less about the number of functions available, and more about mastery:
[…] In practice, bloat comes from the fact […] that Office does so many things that customers just assume the product can do whatever they need it to do. Despite that fact, customers have no idea how to make the product do what they need. This feeling of helplessness that leads to frustration. […]
Bloat is owning a product that you cannot master.
This below is a great observation about the perils of an idea of a “simple mode,” which Sinofsky argues is always a leaky abstraction:
We tried reducing bloat by hiding features […], but that only added to the mystery of the product. Mac, Windows, and Office all went through periods of “simple means fewer” and tried mechanisms such as short menus, simple mode, or adaptive toolbars. But that frustrated or confused people. No one really wanted to use a simple mode and there was always one command missing that was needed, so simple mode became a complicated way to do that one thing that made someone’s work unique.
It was great to see this argument for a broad definition of a bug, as it slides exactly into my post from a while back:
Ages ago in ancient Microsoft history there was a debate on the original apps team about what it means for something to be a bug. Is it a crash? Is it data loss? Is it a typo in an error message and so on? Out of that was created a notion of bug severity, a measure for how serious a bug might be from losing all data all the way to simple cosmetic issues. However, when it came to talking about bugs with product support or ultimately customers the definition of a bug was very simple “a bug is any time the software does not do what a customer expects”. This definition created a discipline of documenting everything reported about the product and always making sure every issue was looked at, even if a code change did not result. The key lesson was how helpful an expansive definition was.
There are also observations and research about how users “debug” the product to make it achieve something they know is possible, but they don’t know how:
We called the futzing document debugging, and it created a frustration that the product was powerful yet overwhelming. People believed a specific result was achievable but getting from point A to B seemed impossible or unlearnable.
And some about the challenges of figuring out what features people use:
[…] Most people didn’t know or care what buttons they clicked on or menus they chose so long as it was working for them—and that meant when asked, “Did you use X?” most people couldn’t recall. To a skeptical press or IT manager (and they all were) that meant unused features.
I should stop quoting and let you read in peace. But, check this out. Lisa wasn’t the only one having linguistic fun:
Early keyboard shortcuts were simple, like using Ins(ert) key to copy text from the scrap (clipboard).
The otherwise excellent note-taking app Bear has an interesting bug that’s worth talking about while it’s still here.
When you’re around to-do items, you can press ⌘. (period) to toggle any task complete or incomplete. It’s actually a really fun shortcut in practice:
But when you have a larger selection with a mixed state (some checkmarks are on, others off), this is what happens:
This feels like an obvious thing to implement, and this is also where the code itself wants to go when left alone.
But this is not great. The rule is: When you have a mixed state, changing it should collapse (or: normalize) the entire selection to one state or the other, rather than perform individual inversion. Try ⌘B in your text editor on partially bolded text, and you can see that collapse in action:
It feels strange to recommend that, particularly as it seems like it loses data. So, what gives?
The first argument is “do not make the user jump through hoops” or maybe “respect a large selection.” If, as a user, I want to actually make sure all my tasks are done, the shortcut not being idempotent means that I now have to go through tasks one by one, and that’s a lot of work – especially since we’re talking about text selection, which is famously unpleasant.
The second reason is that even the UI layer has an opinion here. In the above bolding example, Pages collapsing the selection to bold when you press the B toggle, makes the toggle UI behave exactly as it normally would with a simple selection:
Elsewhere, in Figma, typing a number on top of a “mixed” state changes all the properties of relevant objects to that number:
Imagine how awful it would feel mechanically in both the above examples if your action would still leave the text in the “mixed” state. It would simply appear like the UI broke, since the change didn’t fully “stick” – kind of like those tiny hated moments when you close the door, but it doesn’t latch on and reopens on its own, or when you engage the turn signal stalk, but it refuses to stay put and snaps right back.
There is also one last reason. It’s the simplest one that I sometimes have to remind myself to put in my head before I jump too deep into the mechanics, or details, or technical nuances. Let’s say the toggles invert individually on a large selection. Who would ever benefit from it behaving this way?
In Figma, when choosing a font, you can filter down a list of fonts from “all” to specific categories or e.g. only fonts present in the current file.
But when you type into the search field, the search cuts across all fonts again, ignoring the applied filter. The search effectively lives outside the filter.
In Keyboard Maestro, when adding an action, you can click in the nav to filter down to a specific category. And when you search, the current filter remains active, so you search inside the filter.
Which one is better?
I don’t have a universal rule here, because it will depend both on the UI treatment, and the specific filters and searches people do.
But I think here, my recommendation for Keyboard Maestro here would be to do the same thing as Figma does. I designed that flow in Figma, so I might be biased, but my reasons are:
There aren’t really a lot of options in each category, so you don’t benefit a lot from double filtering.
But the most important thing: For both Figma and Keyboard Maestro, the text field might smell like a text filter and as such expected to be multiplexed with the category filter, but I think this field is actually something else: It’s quick keyboard access, like ⌘F or Spotlight or Raycast. And if you think about it this way, it’s important for it to be deterministic – I can always type “Output Sans,” no matter what state am I in, and get to the font.
On that last note, I find it’s good to look around what you’ve designed once in a while and consider not what the UI set out to be, but what it became – there might be more examples of that around you.
To me, “tap anywhere at the top to scroll to the beginning” is an amazing and underappreciated mobile gesture:
It not only provides an alternative to desktop‘s Home and ⌘↑ keys, but the student laps the teacher here; it’s actually better than every way to scroll to the top on desktop (do you like pressing ⌘↑? do you even have a Home key?), and it’s an icing on a cake of a regular flick to throw the page to the top already being pretty nice.
Tap to return to top is also distinctively mobile in that it allows you to tap just anywhere near the top edge that’s not already a tap target; as far as I can observe, traditional GUIs detest being imprecise in this way, always asking you to click on something specific (although window moving on macOS in the post-title-bar era is also starting to feel similar).
The iPhone gesture seemed to work so well that, over the years, more patterns started borrowing from it. In Bluesky and tons of other apps, you can tap on any tab with scrollable content a second time to scroll all the way to the top. (Again, something that’s hard to imagine on desktop, where you pretty much almost never think of clicking on an already-selected item.)
It’s not just the top, either. In Podcasts, tapping Home goes back to the left:
And in Photos, to the bottom:
To me, the whole “tap to return to the beginning” gesture universe feels ascended to be the core property of the interface. In that way, it is similar to scrolling, undo, copy/paste, arrow keys moving the text cursor, and so on, all inducted to the National Register Of Historic Gestures.
Why? Because these gestures can only blossom if they work consistently, everywhere. You need to start trusting them so much they slide into your subconsciousness. Breaking the gesture in one place will make it less trustworthy in other places, too, ejecting it from motor memory back to the level of deliberate effort, and therefore making it a lot less usable. “Does this thing work here or not?” is a death knell of flow.
The fact that tapping on tabs is idempotent means there’s also no penalty; if you’re already at the beginning but are not sure, tapping it mindlessly won’t hurt or send you back somewhere else.
This is all great. And this is why I’m unhappy Safari started mucking with it.
Safari has tabs at the bottom – starting with two (regular set and “private” set), although you can add more. Above is a long list of site cards, with newest at the bottom. It’s exactly the same situation as in Photos, and yet tapping on either tab doesn’t restore the scroll position. Instead, it opens the settings dialog:
And, tapping around the buttons does nothing.
I would imagine Safari is a pretty important app used by many people, and so this feels like a bad place to introduce an inconsistency that could have a more serious consequences of un-teaching people about tap to scroll to top in the long run.
The funny thing is that the solution is already there: you can tap ··· in the upper left corner to get to the same functionality. The long press on the tab also opens the same menu.
Messing with a “tap to go back to the beginning” system gesture like this means to me the design team doesn’t fully share the understanding of the value of their own creation, or maybe that stewards of the gesture system are not vigilant… or perhaps the awareness is there, but the caretakers aren’t recognized, rewarded, or empowered enough.
It’s similar to the “no, thanks” example I shared before, a possible worrisome tragedy of the UX commons in the making if the respective teams do not change course. Because, wedging that sort of an exception in – even if you have a great set of reasons in the moment – creates a precedent. Inevitably, from my experience, the next team that will want to override scroll to top, or misuse “No, thanks,” will now require less of a justification.
Software engineer Ajitem Sahasrabuddhe recently wrote a 6-post series called “Iron Core” about airline ticketing infrastructure. The entire series is probably too software engineer-y for us, but the third part has some interesting info about a particular 1960s user interface called “cryptic mode”:
Cryptic mode was born from a hard constraint: teletype terminals in the 1960s billed by the character transmitted. Every keystroke cost money. A command that took 50 characters instead of 10 cost five times as much. Commands were compressed to the absolute minimum.
The result is a domain-specific language whose syntax was shaped entirely by economics. AN for Availability Next. SS for Sell Segment. NM for Name. ER for End and Retrieve. No vowels wasted. No words spelled out.
Apparently the official name is “native mode,” but it gained its nickname because… well, see for yourself.
Asking the system for “Availability for Next flight” for February 8, from Nagpur to Delhi, is just 13 characters:
AN08FEBNAGDEL
And the system responds in an equally mysterious way:
** AMADEUS AVAILABILITY - AN ** NAG DEL SU 08FEB 0000
1 AI 416 Z9 C9 D9 Y9 B9 NAG DEL 0840 1030 32A 0
2 AI 416 M9 H9 K9 Q9 T9 NAG DEL 0840 1030 32A 0
3 6E 5317 S9 T9 W9 V9 Q9 NAG DEL 0840 0755 32A 0
With time these commands became wrapped inside more approachable interfaces and GUIs. But they exist under the hood and…
Many experienced travel agents still use it today alongside, and sometimes instead of, web-based agent interfaces such as Amadeus Selling Platform Connect. For a trained operator working a booking-heavy workflow, it is faster than the equivalent graphical interface for the same sequence of operations.
Except today, you get to choose. At the beginning, when “online” didn’t imply internet, and registration computers looked like this, you didn’t have a choice: this was the language you had to fluently write and read.
In his latest video, Shelby from Tech Tangents unpacked, installed, and put to use a truly forgotten product: IBM 3119, one of the first consumer flatbed scanners.
The setup was a small nightmare, needing a rare hardware card installed in a specific computer, an ultra-particular combination of two operating systems working in lockstep, and even some careful memory balancing.
Even after all that, a 300dpi page scanner in the late 1980s was still a force to be reckoned with. It’s hard to remember how enormous scanned files were compared to anything else then, even on a black-and-white scanner like this one. The video shows a simple 90-degree image rotation in highest quality requiring over 9 hours, and I believe it.
But deep inside the video, at precisely 19:31, for only ten seconds, something appears that is absolutely worth celebrating. The nascent scanner software has a “curves” feature that allows you to redraw the shades of gray to capture shadows, highlights, and midtones exactly how you want them. Today, the feature would look something like this, with a real-time preview:
There would be absolutely no way to do something like this in the late 1980s, when just rotating an image is an overnight operation, right? And yet:
How was this accomplished? Absolutely brilliantly. Remember the palette swapping technique? Here, the entire screen’s palette is 256 shades of gray. It’s a very particular kind of a linear palette, and so you can easily take that line and… well, turn it into a curve. Since palette swapping happens on the graphic card, it takes as little as one frame of time, allowing for it to react to mouse movements as they happen.
This must have been mind blowing to experience in the moment. Sure, it’s only a preview, and actually applying curves to the image would take many minut—
No. This is a wrong frame of mind. Here’s my hot take: There are moments in software where the preview is more important than the feature following it. That’s because the preview making things faster isn’t just the difference between finishing something sooner or later. It’s a difference of doing something or not doing it at all. Would you even attempt to use curves if each adjustment took minutes or hours, especially in a land without undo?
I love this preview that hints at what the future will be. I like this clever use of extremely limited technology and tight collaboration between engineering and design. It must have been nice to be in the room whenever someone had the flash of insight to use palette swapping this way.
I want to show you something glorious. This is Bear, the note taking app:
There are desktop apps that get flustered if you ⌘+Tab away and back, misplacing focus or closing a dialog box inside. There are iOS apps that fully reset themselves whenever they get swapped out of memory and have to be reloaded.
But Bear, right here, remembers which note you were on, and exactly where you were in that note, even between phone reboots.
Software is transient and malleable, and one of the hard parts is knowing when that’s beneficial and when detrimental. In real life, you can leave a notebook on your desk, open on a certain page, leave a pen pointing to a specific word – and then depart for a two-month trip to Europe. You will find your notebook exactly how you left it. Why shouldn’t software behave this way?
Also, another thought: This is very likely not something users will complain about when broken, or suggest when absent, even if you go out of your way to open yourself for feedback. Just swapping an app out of memory is hard to understand and “repro” (in engineering parlance). There’s a certain design mindset and taste necessary to notice and care, and a certain vision to carry it through.
The lack of direct user feedback doesn’t mean it’s not worth doing. It just means that there are some things that designers and only designers will know how to properly weigh, describe, and prioritize. If you have a few design-minded users that actually send you feedback like this – treasure them. But most likely this will have to come from “inside the house.”
To me, it’s clear that within Shiny Frog (the makers of Bear), there are people who care about this kind of stuff, and leadership that trusts them. Kudos.
This is perhaps my favourite feature in Lightroom. You press ⇧T, you draw a few lines, and presto – your photo is now even:
This is doubly magical to me. The first part is that this is even possible – that you can straighten the photo in both dimensions after the fact, and save for some parallax nuances the viewer won’t know any better.
For decades, this has been the domain of tilt-shift lenses, but if you ever tried to use one, you know how harrowing of an exercise this is. A tilt-shift lens looks more like a medical device and less like a piece of photography equipment:
The “obvious” way to emulate a tilt-shift lens in software is a bunch of sliders, and Lightroom has those also…
…but that’s still pretty cumbersome in practice, abstracted in a strange ways, like piloting a plane by pulling the linkages connected the flying surfaces: you will admire someone who can do that, but won’t ever want to do it yourself.
Hence the second magical moment: The team created the new interface I showed at the beginning, where you point to things that should be straight directly, and the necessary tilt-shift calculations happen behind the scenes.
Alas, Lightroom didn’t fully stick the landing. The interface is a bit jittery, and missing nice transitions that could help understand what’s going on. But what brought me here was this unpleasant interaction:
What’s wrong with it? If you want to play along, stop here and ponder: How would you improve it? Because this is a classic UI exercise where there are symptoms, and there are problems, and there are principles under the hood of it all.
The first possible improvement: Don’t do a dialog like this. These are ancient and so annoying. Every time I see a centered dialog covering everything, popping up in response to a delicate mouse operation, I want to shout “read the room!” It’s better to drop a little tooltip next to the cursor that automatically disappears: more modern, and more “compatible” with mousing.
Then: Why am I allowed to start and finish an action that the machine already knows won’t go anywhere? Disable the drawing option, put a little “verboten” icon on the mouse pointer, or do something else that will prevent me from drawing a line to begin with.
But that brings us to point three, and how I would approach this as a designer. Because I would – counterintuitively – go the other way and allow the user to draw as many lines as they wanted, and just didn’t permit to commit the entire operation if there were more than four lines on the screen.
Why is that?
It’s the same principle as you see in all the social media composing fields, and in well-trained forms: do not constrain the editing process.
This field is limited to 300 characters, but it’s clever enough to only enforce its limits when you try to post. There is no downside to allowing you more room in the editing process. Maybe you write by constructing a few sentences first and only then combining them into one, maybe you want to see two riffs one below the other to choose the better one, or maybe – this is most likely – you’re not even paying attention and your motor memory is doing the editing for you, instinctively. Use any text editor for just a few months, and cut, copy, and paste, word swapping, and splitting sentences become second-nature gestures – that is, until the UI starts throwing in some arbitrary barriers.
Above in Lightroom, it might actually be easier for me to draw a fifth line and then delete a previous one, instead of doing it in the precise order Lightroom desires, or by dragging an existing line to move it instead of creating a new one.
Maybe an overarching principle would be this: If you are aiming to build something so delightfully direct manipulation as Lightroom did here, you have to fully commit to that stance, even deep in the weeds. Because every time I see a 1990s dialog appear when my fingers are flying fast, I feel like this:
To follow up from yesterday’s post, in Figma, object selection actually goes onto the undo stack. This is because in a professional tool with objects in multiple levels of hierarchy, it might take a while to construct a selection to work on – and since selection is always just one accidental click away from being completely cleared, undoable selection is extra protection.
However, at the same time renaming a file – or changing settings like file access – is not undoable. This is in part because we didn’t feel people would understand they could cancel out their rename this way (Safari too used to have “reopen last tab” under ⌘Z, until it reverted to Chrome’s ⌘⇧T), but mostly because you could accidentally undo through a file rename during regular work if you were not careful, without noticing, and that felt like it’d have more profound consequences.
In some ways, it helped me to think of these not as “ineligible for undo” but rather “living outside of time.” The moment a file is renamed, it will always have been named that way. (For the purposes of undo, at least. You can acknowledge anything you want on the version history screen.)
I’m not saying these are universally correct choices – as a matter of fact, some users find undoable selection (at least initially) pretty confusing! – but mostly sharing these as examples of intentional thinking about what deserves undo, and what should be exempt from it and taken care of elsewhere.
I am a huge fan of all sorts of “recent” features in software; I think they’re extremely helpful in removing tedium, and thoroughly undervalued. A lot of our work is repetitive, even if it’s sad to admit.
My bank’s website not only shows me the last payment I made, but also allows me to click to use the same number again:
2.
The app Transit has a nice list of recent destinations just below the main options:
3.
Google Maps promotes recently tapped-on items to be more visible than they would normally be:
4.
CleanShot X offers something I have always wanted from built-in macOS screenshotting – being able to capture with one keystroke the same area as I delineated last time:
5.
Google Pixel allows you to swap the current wallpaper and three previously chosen wallpapers easily:
What unifies all of these is that “recent” doesn’t live in a submenu somewhere, treated as a second-tier pathway. No, in all of these “recent” is embedded in the fabric of normal interactions, side by side with forward-facing options. I believe this is necessary for any sort of feature like this to be truly successful.
That last Google Pixel example also shows that “recent” isn’t only for repeating something faster – here, it becomes more of a “soft setting,” without introducing a lot more complex UI and interactions that a “real” setting might require.
The gist of it is simple: the mechanics of following a link are not important, and should be replaced by something that can make the link stand on its own. This is important for screen readers, but also for basic scannability: a “click here” label has a lousy scent and requires you to take in the surroundings to understand what it really does. The rule is, in effect, a variant of “show, don’t tell.”
(In modern days, you can also add another transgression: on touch devices one cannot click, but only tap.)
There is a similar rule about button copy design. Button labels, too, should be self-sustainable. Below is a good example (just reading the button lets me understand what I’ll achieve by clicking it), juxtaposed with the bad one (“OK” is so generic you have to read the rest of the window).
Earlier this week, I was passing some train cars on my coffee walk, and saw this bit of UI:
Why are these okay, and “click here” is not? Here’s why, I think: Yes, the ultimate goal is to move a train car, or empty it, or send it on its way. But here, the mechanics matter, too. They’re dangerous. They require preparations. No one says “I’m going to open my laptop and start clicking on links,” but I imagine people say “we have to jack this car” or “we need to lift it.” Even “here” has depth: these are specific tool mounting points. Choosing the wrong “here” will have consequences.
But, going back to the web, avoiding “click here” in strings isn’t always easy. Imagine trying to put a link in the sentence “To change your avatar, visit the profile page.” I’m personally never sure how to linkify it well:
To change your avatar, visit the profile page.
To change your avatar, visit the profile page. To change your avatar, visit the profile page.
Linking “change your avatar” seems correct since it points to the eventual outcome, but then it leaves the actual destination dangling and unlinked – like putting an accent on a wrong syllable. “Visit the profile page” is better than “click here,” but it’s still not scannable. Linking the entire sentence seems strange and complicated to me, and I also disagree with Tim Berners-Lee, who on the page I liked to above seems to suggest this should be…
To change your avatar, visit the profile page.
…just because this might make a user think there are two separate destinations and actions, and contribute a wrong mental model.
You could, of course, simplify this to “Change your avatar,” but while that would work in a UI string, it wouldn’t within a larger paragraph of text, or a blog post.
Wakamaifondue is a web tool to inspect font contents, and it starts by you dropping a font file (.ttf, .otf, or .woff) into a browser.
It handles file dropping so thoughtfully, it’s worth pausing and recognizing it:
Here’s what’s great about it:
You can drop the file anywhere. There is no designated small drop area like in some other apps; every last pixel of the window is ready to receive your file, so you can drop without worrying.
You get a hover state confirming you are safe to drop.
You can drop the file on other screens, too!
Why is all this important? Because dropping a file into a browser is a notoriously frustrating experience. If the tab doesn’t claim the file, left to its own devices the browser will do anything from replacing the current tab with the contents of the file, through opening a new tab, to… starting to download the file you just dropped and ask you for its new location!
It is frustrating when a failure mode of an action is not just that action failing – already here, repeating a drag is more work than e.g. repeating a keystroke – but also you having to do extra clean-up steps.
Wakamaifondue gets this right, and allowing to drop a file on any screen in particular is very thoughtful. Your cursor holding a file indicates your intentions rather strongly – when you see a person wearing a wedding dress, you don’t think “I wonder what they’re up to today?” – so there should be no need to switch to a certain mode or to navigate to an “import screen” beforehand.
Right next to the generic function to delete photos by going through them one by one, my camera has a specific version – Delete All With This Date:
Below the actions to close the tab, and close all other tabs, Chrome has a specific version called Close Tabs To The Right:
In After Effects, next to typical save options, there is this – Increment And Save – which saves a file and changes the number at the end to be one notch higher (Project 2 → Project 3, and so on):
I’m mildly fascinated by these strangely specific accelerators.
The one in the camera is genuinely useful. Photo projects are often day-long affairs where you download the photos at the end of workday, but might still keep them on the card just in case. Allowing to quickly delete a day’s worth of photos makes a lot of sense, saving you from having to go through them one by one in an interface not suited for that kind of operation.
Chrome’s “Close Tabs to the Right” takes a bit of figuring out, but I believe it’s meant to make it easy to clean up after a fruitful research session where you kept ⌘-clicking and opening tabs to learn more, and those tabs now fulfilled their purpose. (Curiously, Firefox also has “Close Tabs To Left” which I don’t understand.)
After Effects’s “Increment and Save” is… I don’t know. Maybe it’s cheap? Maybe it’s honest? A proper version history would be nicer, but that’s a tall order. This is simple and, most importantly, reliable. I still often do the “poor man’s version control” elsewhere…
…so this works for me.
It’s always interesting to me to think whether these kinds of oddly-specific examples are nice gestures toward the user, or treating symptoms in lieu of fixing actual problems. Either way, I don’t think an interface can survive too many of these, as their obscurity and weirdness add up and can contaminate the entire UI.
Would love if you sent me more of these kinds of commands from the apps you use!
One of the casualties of Apple’s otherwise brilliantly executed transition to retina pixels has been the mouse pointer, which remains aligned to what “traditional pixels” used to be, rather than the retina/physical/smaller pixels.
Turn on the zoom gesture from a few weeks ago, and you can see the challenge. The gridlines are ½ logical pixel and 1 physical pixel wide:
This limitation is inherited by most tools: Photoshop, Affinity, xScope, even the built-in Digital Color Meter. It’s not the end of the world, of course, but it can be maddening if you are trying to sample a color from a “half pixel” and the cursor stubbornly skips it no matter how delicately you move. Here it is in Figma:
Of the few tools I tested, only Pixelmator allows to sample at the correct, precise level:
I was curious how would a truly precise cursor feel in general – would there be any disadvantages? – so I built a little simulator that allows a regular arrow cursor to be aligned to “half pixels” or “retina pixels.”
In the process, I discovered that both Chrome and Firefox already receive sub-traditional-pixel measurements for mousing events, so this was even easier to build than I expected. Now, precise targeting in Chrome and Firefox becomes possible:
I don’t personally see any big difference in terms of either upsides or downsides, and I’m curious if you do. iPadOS and its Safari already seem to support the precise mouse pointer, too. That makes me curious: why isn’t it available in macOS? I imagine you could even turn it on by default for apps – or, if you want to be more conservative, make it opt-in.
Pixelmator also shows that the apps can do it without waiting for macOS as the data is already there; they would just need to render the cursor on their own with more precision.
It’s a fun (as always) watch, but as a UX designer, it’s also interesting to try to figure out what are the underpinnings of the things Staecker lists as strange from today’s perspective.
I believe that “CE/T” (clearing and totaling) coexisting on one key is a nod to professional accounting use of adding machines where you wouldn’t want to accidentally enter something into the record twice – so totaling also automatically resets the value and prevents you from making a mistake.
I also believe the strange [+=] rule is only because the keypad has to look forward at the same time it is looking back: it needs to serve as a universal computer keypad where [+] and [=] are separate key, but it also needs to pretend to be an adding machine where one key served both purposes.
(You can spot that the back of the box just allows you to swap the [+] key to be something else.)
Overall, the video is a fascinating tale of an “in-betweener” product that was stuck not just in the middle of a transition from physical devices into apps, but also at the intersection of calculators and adding machines (once two very different lines of products), themselves trying to learn from each other. It also serves as a great reminder that skeuomorphism is not just about visuals and sounds, but also behaviours: tearing off the tape, details of specific keys, nuances of rounding.
It’s not a thing of the past, either. In my post about determinism I linked to Apple’s recent travails with the deterministic Clear button (part one, two, and three). A few years ago, Apple also changed the built-in iPhone calculator from its “desktop calculator” roots to a more modern model where you get to input the entire equation before you see the result. But that change had bigger consequences; for example the [=] key could no longer repeat an addition. People complained, and Apple added it back – but the change feels incompatible with the new system and potentially confusing:
Elsewhere, the entire iPhone is an in-betweener, as the keypad coming from calculators is incompatible with the keypad coming from phones.
At this point it seems the calculator keypad will win, but transition has been over a century in the making. Staecker’s video is a good reminder how important, but also hard it is when you try to make these transitions happen faster.
A few years ago, I suggested adding a new interaction to Figma. If your text cursor was on a misspelled word (anywhere inside, or the edges), you could press Tab to quickly accept the suggested correction, without even seeing it:
Independently, Google Docs approached it from a slightly different angle, but landing on a similar interaction – in their version there’s a small visual callout, although you can still press Tab (and then Enter) to accept the suggestion:
I know the Tab key has a lot of jobs – from indenting bullet points to jumping through GUI elements – but in this context this new addition doesn’t seem to be in conflict.
(Should I write a long photoessay about the Tab key, similar to the ones I wrote for Return/Enter and Fn keys?)
Since we added it, I’ve really loved how it feels. From various typeaheads and autocompletes elsewhere, Tab has a strong “forward movement” energy so it makes conceptual sense, and it’s just really fun to go around and quickly fix your writing this way.
I think a lot about how to make keyboard interactions feel superpower-y: a good keyboard shortcut on a large key, a tight interaction, a blink-of-an-eye velocity – something that’s eminently designed to lodge itself in your motor memory as quickly as possible, as it builds on top of prior motor memory. I’m biased, of course, but I like the “no scope” Figma version more, and it has that feeling to me.
Recently, spelunking in the preferences of Photoshop 2025, I found this extremely curious thing:
To transcribe:
Focus mode limits the appearance of certain optional user interface messages so that you can use Photoshop with fewer interruptions.
With this option enabled:
The Welcome screen will not include “what’s new” feature descriptions
Blue in-product alerts promoting discovery and use of certain features will be suppressed
What’s New will not auto start when Photoshop is launched
The color mode preference will be auto set to “Neutral Color Mode”
The three first options should be self explanatory. Neutral Color Mode is sort of the “graphite” option of Photoshop’s UI where the (already rare?) accented blue elements become white instead.
As much as I’ll always applaud a piece of software working on annoying you less, this is all so very strange. I don’t mean that the last option seems unrelated, and the first and third one kind of mutually exclusive… but just the very idea of shoving it in as an opt-in in the last tab of settings, under “technology previews”, and asking people for feedback feels peculiar to me.
Not to spoil the outcome, but even this “technology preview” is completely gone in the updated Photoshop 2026. I wonder if this is fallout from a mangled launch (even for those few who I imagined turned it on, the option didn’t live up to its promise), but also perhaps a political fight inside Adobe between product and growth teams? I bet we’ll never know.
I do not personally have a grand unified theory of how to explain things or announce features in products because it’s so situational, and I understand that especially Photoshop given its age might be the hardest difficulty level. I’d personally prefer to receive announcements of new features over email so I can read them at my leisure, and with each new thing or change linked to a playground that would allow me to experience it in the best way – but I can’t say with any certainty that this would work for everyone.
But I would expect people on the Photoshop team to have more experience here, and this focus mode approach just feels a bit… naïve to me. My two warm takes: 1. People aren’t generally as frustrated with how features are announced, but with what features are. 2. Why wouldn’t everyone deserve the gift of focus?
⌘T is a very important shortcut in Slack. It allows you to quickly talk to someone just by typing in their name. I use it probably dozens, if not hundreds of times a day.
⌘T is right next to ⌘R, which reloads Slack. Occasionally, on the way to ⌘T, my fingers graze ⌘R. Fingers being fingers, I immediately realize something went wrong and wince, and within a second or two I witness Slack completely reloading. It’s not a big deal – no data is lost, and the reload is only 5 to 10 seconds, but when you move fast, it feels like eternity.
⌘O is a very important shortcut in Finder. It opens the selected file in the correct app. I use it probably dozens, if not hundreds of times a day.
⌘O is right next to ⌘P, which prints the file I’m pointing to. Curiously, and in contrast with most apps, the print function is not gated in any way by a confirmation dialog box, or an intermediate print settings window.
So, occasionally, on the way to ⌘O, my fingers graze ⌘P. Fingers being fingers, I immediately realize something went wrong and wince, and within a few seconds, the lights in my old apartment dim for a second. Then, far away, I hear the recognizable sound of my laser printer spitting out a page.
Gamers used to deride Windows key for automatically ejecting them from the game to the desktop, before an option to disable it started appearing in gaming keyboards. (Some of the professional gaming leagues were very strict about how a player could use their keyboard.)
Similarly, professional Excel champions and players started physically removing keys: In Excel, F1 (right next to an often-used F2) opens the help dialog and slows you down.
I served as a judge for the ModelOff Financial Modeling Championships in NYC twice. On my first visit, I was watching contestant Martijn Reekers work in Excel. He was constantly pressing F2 and Esc with his left hand. His right hand was on the arrow keys, swiftly moving from cell to cell. F2 puts the cell in Edit mode so you can see the formula in the cell. Esc exits Edit mode and shows you the number. Martijn would press F2 and Esc at least three times every second.
But here is the funny part: What dangerous key is between F2 and Esc? F1.
If you accidentally press F1, you will have a 10-second delay while Excel loads online Help. If you are analyzing three cells a second, a 10-second delay would be a disaster. You might as well go to lunch. So, Martijn had pried the F1 key from his keyboard so he would never accidentally press it.
I enjoyed this essay that presents prying off the key as a rite of passage:
Removing the F1 key from the equation is just the beginning. By embracing the keyboard-centric approach, you have the opportunity to become an Excel Wizard!! Okay, maybe that’s not a technical term, but it perfectly captures the essence of those who navigate Excel solely using the keyboard.
And I particularly liked this tongue-in-cheek answer telling people they could construct their own homemade molly guard to protect against “fat-fingering”:
Here’s an alternative snippet that can be used:
Use bits of plastic or cardboard to make a tiny box that fits around your F1 key.
Affix this box with duct tape, so that the F1 key is guarded.
Fool-proof, works on any key, and can easily be reversed if needed!
Obviously, none of this can help me with my ⌘R and ⌘P woes, so, two final thoughts:
If your app has a well-trafficked shortcut, it’s worth thinking of the shortcuts immediately adjacent to that one. Could they cause any inadvertent damage or confusion?
Apps and operating systems should very easily allow you to unset a keyboard shortcut, in addition to setting or changing it. (Unfortunately, this is not as common as it should be.)
This video from Marblr about adding fall damage to Overwatch is really intense – 45 minutes of length and a lot of footage of frantic gameplay – but really informative, too.
It’s a great case study of how something seemingly really simple – deducting health from the player as they fall from height – can be a complicated thing to figure out in all the detail.
I never played Overwatch and rarely play videogames anymore, but many of the lessons here more universal for any sort of UI and system design:
You will have to introduce tactical inconsistencies for the system to feel consistent, but be careful as there might be a point those inconsistencies start to outweigh the whole thing.
Wanna learn how you and others feel about something? Overcrank it to make the feelings come out more easily. (And to find bugs.)
There will always be tensions between what the data says and how you feel about something. (I was surprised how often the word “intuitive” entered the picture.)
Also, it’s just a really well-made video, filled with little presentation and storytelling details that elevate it. I wish more videos like this existed for UI mechanics.
But maybe the most important takeway? You don’t have to choose between rigor and fun. You can have both.
I just stumbled upon a nice little power-user innovation in Chrome’s Web Inspector.
In Safari, and previously in Chrome, when editing CSS properties, you’d get a usual editing typeahead for the property name, and then the same on the other side for the property value.
In newer versions of Chrome, the typeahead menu works as before on the right side. However, the menu on the left side also includes the right side.
I think this is really clever in this context – not just to speed you up, but also to aid understanding. Just like the inert mouse up and down in the previous post could serve as a safe “peek” into the values, this new interaction can quickly allow you to explore the CSS space if you are curious, or if you only lightly remember part of the name, or even just one of the values.
This blog is authored in Apple Notes, and some time ago Notes added quick linking via typing >>, and that has a similar effect: The interactions are so nimble and precise that it is very easy to link to something, but a nice side effect is that it also feels very welcoming just to type a few letters to remind yourself of a title of an article, and then cancel out.
The downside of the Chrome change is, well, more stuff matching, but I think the audience for this UI is going to be okay with that.
I know we’re probably collectively a bit tired talking about macOS Tahoe, but I just noticed something that I think is a good example of how small details can ladder up to bigger things.
This is macOS Sequoia (the pre-Tahoe release) and a typical pop-up button:
One clever thing macOS has been doing since basically the dawn of GUIs is that upon clicking on a button like this, the currently selected row will be in the same place as before you clicked. (As opposed to, for example, the entire menu appearing below like it would from a top menu bar.)
This has interesting and often underappreciated consequences. It allows you to orient yourself quicker since you don’t have to find the selected option again. And, it saves you movement overall: the next or previous option will always be at the absolutely shortest possible distance. (Of course, the approach also has some challenges,for example if the button is positioned close to the top or bottom of the screen.)
There’s another clever thing that happens throughout macOS: All the menus work using a classic click-to-open and click-to-select sequence, but they are also usable via the slightly more advanced, but faster mousedown-drag-mouseup gesture.
These building blocks work together and mean that selecting the next option can be as simple as a little flick of a mouse.
Now, check out macOS Tahoe (current release):
You will notice that iCloud Drive, upon clicking, is now misaligned both horizontally and vertically.
On the surface, this feels just like a visual blemish – slighly embarrassing, but without much consequence. But check out what happens if you hold your mouse button at a certain position, and then release it without moving:
The stability of macOS’s interface and the thoughtful set of aforementioned rules allowed for an emergent fast behaviour: mouse down and up meant you could “peek” into a menu safely, or you could change your mind right after seeing what’s inside. In a bigger sense, it created a certain trust between you and the operating system: it’s worth learning those gestures, as they will be rewarded.
In Tahoe, some of that learned behaviour – by the way, I see it in all of these buttons, not just this one – will now work against you. Now, you can accidentally change an option without intending to do so.
Is it a big deal? No, not really. This likely – hopefully! – simply fell through the cracks in a rush to get Liquid Glass out the door, rather than no one being there to care, or no one understanding that all these gestures add up in aggregate, creating a GUI that feels fast, trustworthy, and catering to your motor memory in a way that elevates your experiences with the interface in the long run.
But I’d feel better if it wasn’t almost half a year since the release, and if we hadn’t already seen other things exactly like it.
Mac allows you to assign keyboard shorcuts to menu items, but the interface is clunky – you have to select the app even if you just came from it, and then type in the menu item name by hand without any assistance:
Other tools, like Keyboard Maestro, do something similar. You either have to type it again, or you can point to it, but in a replica of the menu of the app shown in a very different style and orientation:
But this week I learned of another app, KeyCue, that approaches this differently. You simply point to the menu item and hold the desired key for a while:
Okay, this is not a universal endorsement. The feature works clunkily, and KeyCue as a whole is way too comfortable adding itself to login items without asking.
But as far as singular interactions go, this is great and eye-opening. It made me realize that the previous things I’ve shown – System Settings, Keyboard Maestro – are really not GUIs, and they don’t practice direct manipulation. They’re still partially command line interfaces dressed up in GUI clothing.
We kind of lightly made fun of Jony Ive going angelic on “staying true to the material” and things being “beautifully, unapologetically plastic.” And there is, of course, value in command line and those kinds of approaches. But this part of KeyCue at least is unapologetically a graphical user interface, and it is nice to still be surprised in this space.
I keep thinking about this very good 11-minute Not Just Bikes video about traffic calming. In it, a simple argument is made: the posted speed limit of any given street or road doesn’t really matter. What matters is how the street feels. Generously wide and separated lanes, sparse traffic lights, and the road being straight past the horizon will make you unconsciously speed up. Reducing the posted speed limit or adding flashing YOUR SPEED signs won’t help:
The truth is that many drivers will not slow down because of signs or speed limits. They’ll slow down either because they don’t feel safe, or because they’re afraid of damaging their car.
The only answer is redesigning the street for the desired speed limit – narrowing the lanes or joining them, creating choke points and speed bumps, adding posts and planting trees close to the road, and even adding visual cues like “dragon’s teeth.”
One of the great thing about driving in the Netherlands is that it’s rarely necessary to look at the speed limit. The road design takes care of that for you.
There is an app I use a lot called Forklift, a suped up Finder, with one of its functions being syncing files to a remote server.
In its version 3, the syncing window looked like this:
This is a pretty straightforward and dependable function – and I’ve depended on it for years.
I recently updated to version 4 to check it out, particularly since it promised faster syncing. But I was thrown aback by how it randomly deteriorated:
It’s not that there seem to be some UI challenges: the new icons make it harder to understand hierarchy, and one of the switches starts with “Don’t” in contravence of rules of avoiding double negatives.
No, the worst part is this:
This is a new temporary state that meant to help me understand the details of what’s changing.
On the surface, it’s a thoughtful thing. But it’s done in the worst possible way for this kind of a power-user interface: It’s very slow to invoke and slow to cancel. I often activate it by accident – it makes large swaths of UI a minefield where you can no longer rest your cursor safely. It also changes the hierarchy of the output in a way that’s confusing – and it even animates the text wrapping in a distracting way. Then, if you press Esc instinctively to get rid of whatever happens, the window closes altogether.
It’s a “delightful,” luscious transition that is completely out of place. I think this is how many people misunderstand craft – that it’s only about “high polish” without any thought underneath. Here, the effort was spent on executing something that couldn’t be saved this way and needed a more serious rethink. It seems like its creators forgot who’s using the app and for what, and embarked on accidental UI calming.
There are other challenges along the same lines, both downgrades from version 3:
when the app analyzes the differences, I can no longer press the Sync button and walk away
even when the button becomes active, I can no longer press Enter to activate it – I have to use the mouse
In version 3, I could invoke Sync, immediately press Enter, and get on my merry way, with syncing continuing in the background. It was exactly what I wanted. Version 4 slows me down by requiring me to pay constant attention to the interface: it matters where I rest my mouse, it matters when I click the button, it matters what input device I use to commit.
It’s okay to think of friction and sometimes transitions are indeed very helpful for UI calming to avoid drastic movements or accidental activations. But here, this isn’t great at all; the creators of Forklift promised me faster syncing and achieved the opposite.
I have been enthralled with this tiny feature in Google Sheets called “Show edit history,” which premiered in 2019:
Mind you, it’s not unconditional love. The execution feels a bit clunky, showing the edit values in a pop-up rather than in situ, with formatting that feels too heavy, and an awkward “No more edit history” state rather than just disabling the button.
But! Just its very presence here is delightful. Version history is often this huge, comprehensive, perhaps disorienting mode you enter that by design deals with the entire file. It always feels like a longer trip:
But edit history reimagines the feature from the perspective of the cell. You can just peek inside, quickly and effortlessly. Right click menu, a few arrows, I learned what I needed, and I barely even moved my hand. It’s a perfect example of the rule “to make something feel faster, make it smaller.” It’s like picking your newspaper at your doorstep in your pajamas rather than having to dress up to go to the newspaper store.
(…he said, dating himself and perhaps also thinking of The Sopranos for some reason.)
This kind of reimagining of something that already exists (see: undo send in Gmail) can be really hard, and I don’t even imagine Google Sheets was the first with this idea – but for me seeing this remix was eye-opening, and it inspires me to this day.
One of the ways I like to do development is to build something, click around a ton, make tweaks, click around more, more tweaks, more clicks, etc., until I finally consider it done.
The clicking around a ton is the important part. If it’s a page transition, that means going back and forth a ton. Click, back button. Click, right-click context menu, “Back”. Click, in-app navigation to go back (if there is one). Click, keyboard shortcut to go back. Over and over and over. You get the idea.
It’s kind of a QA tactic in a sense, just click around and try to break stuff. But I like to think of it as being more akin to woodworking. You have a plank of wood and you run it through the belt sander to get all the big, coarse stuff smoothed down. Then you pull out the hand sander, sand a spot, run your hand over it, feel for splinters, sand it some more, over and over until you’re satisfied with the result.
This is a clever metaphor and I wish I thought of this before. What follows is a specific story of finding a few dead pixels in between related interface elements, which is an absolutely perfect example of something with non-linear frustration: It might not register at all on the first try, but it will bother you 1,000-fold on the 20th go.
I was just on Internet Archive earlier today, uploading some documents I scanned this weekend. Their UI is… how would I put this… let’s just say Internet Archive makes Teams feel like Linear. (I love Internet Archive and their work and mission, but let’s be honest here.)
Yet, I found something marvelous. Whoever put the upload form UI together knew there will be people like me who’ll be filling out 20 of these forms one right after another. So they made sure every pixel in their form is clickable to edit the nearest field. And I mean, every pixel.
Whoever you are, you have my nod of recognition. In at least this one respect, it’s clear someone spent a lot of time with the sander.
This is of course competence porn, made even better by the dry Polish lektor-like delivery. But it’s also a puzzle. I watched this so many times. There are so many great UI lessons in here:
You can absolutely put graphics inside a textbox
Sparklines rule
Slider is still the best UI element in history
Previews don’t have to feel like training wheels
Synchronizing sounds to visuals is so powerful (see: turn signals on a car dashboard)
I found myself thinking about how you’d design something that feels real-time, but also needs to be resilient against typos, and has a distinct “commit” moment (which is what I think those yellow flashes are); some of the best moments in the video are the quick fixes that aren’t narrated.
Ultimately, this also shows how powerful and underrated plain text can be as interface. It’s a bit like designing straight in CSS, operating at the weird intersection of motor memory, creativity, and abstraction. (Is there a CSS editor that feels more like this?)
On top of all of this, the act of building the track this way is also how the finished track would sound like. Amazing stuff.
Remember all these jokes that went like this?
[God looking at a pug dog for the first time] What the hell did you humans do with my bad ass wolf I gave you?
Imagine sitting the creators of the typewriter in front of YouTube and having them watch this video.
Let’s say you are in Reeder (an RSS reader for iOS), looking at the list of posts, and already from the title you know you don’t care, and you want to mark it as read.
You can tap to see it and then swipe back the moment it shows. This is the slow path.
There is a faster path. Reeder enables you to slide right or left on the item. You get nice haptic feedback, and many apps support this kind of an interaction.
But there is an even faster path.
You can tap to see it and immediately swipe back. Your thumb is already there on the left anyway, and the distance is a lot shorter now.
Like every advanced gesture this takes a bit of practice, but I noticed I started doing it instinctively, without even thinking.
This happening required two small design details: The original slide transition to be interruptible at any moment, and the app to support swatting/draging the incoming item away even if my finger was nowhere near it. Both are clever, and both feel very welcome, because they enabled this emerging (to me) behaviour that made going through the list snappy without me even realizing.
This might be a good modus operandi: Think of the slow interaction. Think of its fast version. Then, think some more.
Nicely done, Reeder team. (Or, if this is a default iOS behaviour, nicely done, Apple!)
A 16-minute video from Ahoy from last year about Chris Sawyer, creator of Transport Tycoon and Rollercoaster Tycoon games from the late 1990s.
The video focuses more on the economics of the industry and some technical details, but what’s interesting to me was how tight those two games felt in terms of UI. They have a shared custom GUI, they are assembly-coded, and they felt perhaps like the last instance of a graphical user interface where it felt there was nothing standing between you and the pixels.
I know those are games and not productivity apps, but they can be inspiring for those, too. You can download OpenTTD, which is a modern recreation of Transport Tycoon Deluxe that doesn’t require emulation, and it still captures the snappy and tight feeling very well.
I’m thinking about it in particular because the web took a lot of that away. The web loves latency and loose interactions and reflow and temporary fonts and CSS leaks and text sticking out of the box and many other papercuts. It’s nice to be reminded of the world where things were closer to the metal, and how that felt as a user.
One of my favourite recently-noticed little patterns is this one thoughtful accelerant in iOS Photos.
If you want to add a photo to an album, you normally have to choose from a list of albums:
However, once you do that one time, a new menu option appears. It’s effectively “Add again quickly to the album you just chose” (Fiałka is the name of my cat):
That skips the album selection altogether. It’s always only just one album you used more recently, so it’s relatively simple… but so helpful. You often, after all, want to add more stuff to the same album, and it saves you choosing the same album over and over again.
This is great because it flattens the option space to zero options, which mirrors how we all think when we’re focused. It’s tunnel vision exactly when you want it.
I have always been a fan of both “repeat”-type actions and smart “recent”s, and consider them a truly underappreciated secret weapon. Those little savings really add up over time – in saved time, in less tedium, and in avoided mistakes. (Imagine not only having to choose the same album for 30th time in a row, but also… making a mistake doing that and tapping on a wrong one! Then the frustration very quickly compounds, as you have to recover from something that felt completely avoidable.)
I always respect designers of interfaces that invest in functions like these. There is also an anti-corollary to this, which is: if there’s only one option, consider not even asking. Slack seems to excel (derogatory) here:
The second one is somewhat defensible since it’s a settings dialog you enter at your own will, although the active “Re-generate answer” when I haven’t done anything (and nothing can be done) feels overbuilt.
But the first of these always appears on a way to other settings (like adding emoji), and it’s even worse than the Remember me? examples because it repeatedly stops you for absolutely no reason at all.
One of the frustrating patterns for me is a dialog box that doesn’t offer “skip it next time” option, or even just defaults to remembering.
My go-to examples? Apple’s Remote Desktop which always throws this thing up on connection:
And this in Photoshop upon saving a PNG file, which has been there forever:
I never change these options. These are flow-killers; trees have grown to maturity as I have spent collective hours in those dialogs over the years/decades, even though they serve no purpose for me.
(The worst part might be if you forget this dialog waits, and move on to do other things, and the operation you thought was completed never actually finishes.)
When I first learned about this book from Jacob Geller’s video just months ago, I thought this was another example in the vein of The Power Broker – a perfectly Marcin-coded book that somehow escaped me knowing about it for decades.
“Pilgrim” is from 1983, and is a story of a pianist discovering the classic videogame Breakout, and trying to perfect his own gameplay.
I love so many stories of videogame mastery, because at times they feel the closest we got to Doug Engelbart’s dream of incredibly effective machine operation somewhere deep below the threshold of consciousness: You and the computer becoming one, eyes and fingers forming feedback loops so perfect they cease to be noticeable.
Here I am alone in a pitch-black hotel room, a middle-aged man with some time to kill, getting ready to check out some jazz clubs in Greenwich Village, in possession of an early cretinous offering from a gold rush grab bag of tuby thingies coming our way from hundreds of decision-making puzzle peddlers throughout the new electric “entertainment” industry. And now instead of playing the game it‘s packaged up to be, I‘ve gotten into more or less occupying myself by outlining invisible triangles across the screen of a TV doodling machine. What am I doing?
Unfortunately, as you can maybe already sense, the book is an overwritten, ponderous, and pretentious mess. “Beach reading, it ain’t,” quipped a Kill Screen reviewer in 2013. But there are some interesting parts in it.
Before, the piano was the quintessential human instrument. Of all things exterior to the body, in its every detail it most enables our digital capacities to sequence delicate actions. Pushing the hand to its anatomical limit, it forces the development of strength and independence of movement for fourth and fifth fingers, for no other tool or task so deeply needed. This piano invites hands to fully live up to the huge amount of brain matter with which they participate, more there for them than any other body part. At this gnetically predestined instrument we thoroughly encircle ourselves within the finest capabilities of the organ.
Then a typewriter, speeding the process whereby speech becomes visible, the extraordinary keyboard for sequencing and articulating perhaps awaiting a still truer sounding board, strings, and tuning, a still more suited canvas for thought.
Then TV.
This arrives at page 26. Alas, it’s kind of downhill from here.
The author visits Atari (imagine that!) to learn that the programmer of Breakout doesn’t really understand what makes Breakout so alluring. The game perhaps lucked in to being so imminently playable, and then replayable.
I’m interested in designing for mastery. We should not rely on luck that separated a classic like Breakoutfrom a hundred other games from that era that felt awful to play and were immediately forgotten.
Sure, Sudnow definitely takes Breakout way too seriously:
Maybe I can remember the five shots by putting pieces of tape on the TV cabinet to mark each paddle destination, I say to myself, even though it seems that would undercut true learning. It’s bad practice to learn the piano by writing the names of the notes on the keys, much better not to use a code, to grasp the layout of things by their own looks and feel. And I can’t carry Scotch tape to a Breakout tournament.
But in a way: why wouldn’t you?
In fact it’s already happening. I’ve found myself playing with the cursor on my word processor just for the hell of it, seeing if I could track it across screen and get it to stop at every comma in the text.
The word processor (or any other app you use often) operating at the speed of fingers unlocks superpowers, and then some.
There’s one experience in particular at the word processor that gets me downright angry at times. There’s no more of that room for finger breathing while you awaited a carriage’s return. You reach the end of a processed line of text and if your word becomes too long for the margin while there’s still alloted space to get it underway, it splits in the midst of your articulation and your voice instantaneously reappears six inches to the left, a quarter of an inch lower. The computer can’t know what you’re about to write, not yet, not a word or even a letter in advance, has to wait and merely calculate how things are going in order to then “decide” where to put the sound. ¶ Before, you felt a big word welling up, hit the carriage return, lifted off from the keyboard just a bit, reorganized your grasp, and dug back into the improvisation with a renewed rhythmic mobilization to continue. And some of the things you found to say, you found because you said them that way.
This was a fascinating tidbit, this reflection on how small interactions can change the nature of creative process.
If this book was cut to 20% of its size, those fascinating tidbits would stand out more, and the book would still be of value today.
But despite this complaint, I miss people writing about using computers this way. Such a big chunk of my struggle with computers today is fighting with it because I expect a better connection between my fingers and what’s happening onscreen.
I wish more designers understood how important that is.
Google Maps is dying a tragic, public death by a thousand cuts of slowness. Google has added animations all over Google Maps. They are nice individually, but in aggregate they are very slow. Google Maps used to be a fast, focused tool. It’s now quite bovine. If you push the wrong button, it moos. Clunky, you could say. Overly complex. Unnecessarily layered. Perhaps it’s trying to do too much? To back out of certain modes — directions, for example — a user may have to tap four or five different areas and endure as many slow animations.
Funnily enough, I feel that way about Apple Maps. I abandoned it since small things felt heavy, mired in superfluous swipey animations that felt like driving a 1960s car. Luckily, this was at the time Google Maps redesign its tiles to match Apple’s, so I got what I wanted to begin with, although in a slightly shady way.
I miss Sublime Text and might take it again for a spin (VS Code and Atom felt slow, Nova is delightful but also struggles in performance, even on simple things).