What’s the best way to interact with a computer? Do we want conversations, or do we want to poke it like a thing?

Venkat Govindarajan writes:

John Siracusa in episode 19 of his pre-eminent podcast Hypercritical:

What were the earliest mass-market PC interfaces like? . . . they were like conversations…you’d tell it what to do, it gives information back to you . . . That was the basic paradigm until the Macintosh . . . What it gave you was not a conversation, but a thing . . . you could poke the thing and see how it reacts . . . it worked like a physical object . . .

I [Govindarajan] really liked the comparison of early command-prompt user interfaces to conversations. It struck me that today’s AI assistants (Alexa, Google Assistant, Siri) are all based around having conversations. If these systems ever approached anything close to human intelligence and common-sense, perhaps having a conversation is the best way to interact with AI. But I wonder if there is a better interface to interact with AI? What’s the next leap from conversational AI? Perhaps Augmented Reality is the answer — artificial intelligence dispersed in our lived reality, giving us glancable information and the illusion of a physical object we can interact with. I wonder if this is why Apple is so bullish on AR as well.

Or maybe the best way to interact with artificial intelligence is the same way we interact with other people — using conversations.

12 thoughts on “What’s the best way to interact with a computer? Do we want conversations, or do we want to poke it like a thing?

  1. How geeky is it to describe a bash session as a “conversation?”

    I meant “Bourne shell” when I wrote it, but with the new media having found profit in rage, perhaps conversation have turned more and more into bash sessions.

  2. Yesteryear’s interfaces were about dialogs because a typewriter terminal interface is linear, one-dimensional.
    (A programmer would still see the machine as an object to poke, and a mainframe system throwing input masks at a data entry clerk was entirely in control of this “conversation”.)

    Xerox PARC created the graphical interface, which is two-dimensional and information-rich; it makes the functionality of a program discoverable because it’s all explicitly there in the menus. This is easier to learn and to navigate than the older interfaces that you either could not use without a handbook, or that left all of the control to the computer/programmer.

    The current AI assistants listed all use a speech interface, which like the typewriter interface, is linear, and that is why the interaction flow is superficially similar to that. There’s nothing in “artificial intelligence” methods that requires this, unless you want your AI tool to pass a Turing test.

    (AI algorithms are interacting with us in computer games through tactics and gunfire, not through a conversation. Grammarly looking over our writing does not have a conversation with us. Google maps picking a route for me is not an “object” but is nit having a conversation, either. A chess AI is not having a conversation, it expects a move and produces a move.)

  3. Back in the good old days (boy, do I sound like a grumpy old man now) searches used to look for all of the terms in the search expression. So a search like “traveling salesman R” would get me solutions to the traveling salesman problem written in R.

    Now, google search does fuzzy match, so the top result is likely to be a joke I’ve already heard about the traveling salesman and the farmer’s daughter.

    What a joke!

    • That comment would carry a lot more weight if those google search results had returned a single joke among all the R solutions on the first page if results for me. (I must admit, I started with “travelling salesman” solution R and then simplified this to your version, but still…)

      • Also back in the pagerank days google had enormous issues searching for R related things and you’d get basically zero results on page one. My memory of this is that this improved greatly after 2010ish.

  4. Suppose you had a robot servant that observed you and anticipated your needs, bringing you a drink as you realized you were thirsty, say. The robot would not be having a conversation with you–ideally, you’d give it no (conscious) commands. You wouldn’t see it as an object–you wouldn’t poke it. The reason we talk to or poke AI is that it’s not yet smart enough to do what we want without us telling it or poking it first. A sufficiently intelligent and capable AI wouldn’t interact with us most of the time, it would simply observe and act upon us. In this scenario, the AI is the user and we are the objects, albeit happy objects.

    If that seems weird or creepy, consider that your heart, lungs, digestion, etc. work the same way when they’re working well. You give them few if any conscious commands, and they deal with your needs as you have them. The intelligence of advanced AI, then, is in its ability to behave as if integrated with your nervous system while remaining only an observer of it. It is a deductive and inductive intelligence; discovering our needs with as little intrusion as possible is the primary problem it solves, a problem still well beyond devices like Alexa that just use their intelligence to interpret commands and to execute procedures that fulfill those commands.

  5. I guess in this context, “Conversation”, in contrast to physical object which you can poke at to get what you what, is that during “conversation” you both poke at each other.

    So, by this definition, Google search is a conversation, since Google is constantly poking at you, trying to guess what you need, learn you favourites, log your behaviour etc.

  6. A big fan of Bret Victor’s A Brief Rant on the Future of Interaction Design: https://worrydream.com/ABriefRantOnTheFutureOfInteractionDesign/

    The way we physically interact with (most) technology right now is pretty bad, especially on the phone/tablet end of the spectrum (keyboards and mice still give some haptic feedback). I think the rise of conversational assistants is mostly a response to the input options on portable tech being really bad; it’s easy for speaking to provide a better interaction medium than a flat pane of glass, but it doesn’t mean that speaking is a good way to interact with a machine.

    • From the followup, https://worrydream.com/ABriefRantOnTheFutureOfInteractionDesign/responses.html :

      What about voice?
      Sure. Let’s use voice for the things people use voice for — asking questions and issuing commands. But I’m personally interested in tools for creating and understanding.

      Creating: I have a hard time imagining Monet saying to his canvas, “Give me some water lilies. Make ’em impressionistic.” Or designing a building by telling all the walls where to go. Most artistic and engineering projects (at least, non-language-based ones) can’t just be described. They exist in space, and we manipulate space with our hands.

      Understanding: If you simply want information — “What’s the price of AAPL over the last three years” — then an “oracle” like Wolfram Alpha is fine. But I believe that deep understanding requires active exploration, and I’m much more interested in explorable environments. Look at the interactive graphics in the Ladder of Abstraction essay, especially the later ones. You come to understand the system by pointing to things, adjusting things, moving yourself around the space of possibilities. I don’t know how to point at something with my voice. I don’t know how to skim across multiple dimensions with my voice.

  7. I was wondering if the following was relevant:

    “For example, we may never fully understand the physical world. Nor how people think, interact, create and or decide. In ML, Geoffrey Hinton’s 2018 YouTube drew attention to the fact that people are unable to explain exactly how they decide in general if something is the digit 2 or not. … However, prediction models are just abstractions and we can understand the abstractions created to represent that reality, which is complex and often beyond our direct access. So not being able to understand people is not a valid reason to dismiss desires to understand prediction models.

    In essence, abstractions are diagrams or symbols that can be manipulated, in error-free ways, to discern their implications. Usually referred to as models or assumptions, they are deductive and hence can be understood in and of themselves for simply what they imply. That is, until they become too complex. For instance, triangles on the plane are understood by most, while triangles on the sphere are understood by less. Reality may always be too complex, but models that adequately represent reality for some purpose need not be. Triangles on the plane are for navigation of short distances while on the sphere, for long distances. Emphatically, it is the abstract model that is understood not necessarily the reality it attempts to represent.” https://www.statcan.gc.ca/eng/data-science/network/decision-making

Leave a Reply

Your email address will not be published. Required fields are marked *