Why Can a Beautiful Website Still Feel Hard to Use?
A beautiful website can still feel hard to use because aesthetics shape the first judgment, while usability emerges from completing a goal; when navigation scent, expected placement, feedback, accessibility, or recovery fail, the visual halo fades and the task becomes costly. Summary
The website looks expensive. The type is composed with care. The motion is restrained. The product photography has the calm confidence of a company that has never held a meeting with fluorescent lighting.
Then a customer tries to change an invoice address.
The account menu uses an unfamiliar symbol. “Workspace” contains billing, but “Settings” does not. The save action appears only after the page scrolls. When the form rejects the address, the error message says that something went wrong and removes the entered postal code.
Nothing about the page suddenly becomes ugly. Yet the experience does.
A polished surface can improve the first judgment, but it cannot complete the task on the user's behalf.
Beauty reaches the user before usability can
The timing matters.
A person can judge visual appeal before the interface has had a fair chance to prove anything about use. In experiments by Gitte Lindgaard and colleagues, 50 milliseconds was enough for stable judgments of website appeal. Alexandre Tuch and colleagues tested website screenshots at exposures from 17 to 1,000 milliseconds. Visual complexity and prototypicality affected aesthetic impressions at the shortest exposures.
At 17 milliseconds, nobody has found the return policy. Nobody has used the keyboard. Nobody has misunderstood a label, recovered from an error, or completed a purchase. The page is still only an image.
Beauty can win the race before usability reaches the track
People can form a stable visual verdict while the page is still only an image and no goal has been attempted.Complexity already changes aesthetic judgment
The page is still only an image
Stable visual appeal can already form
Category expectations become easier to read
No task outcome has been observed yet
First impression → Interaction evidence begins later.
Tuch et al., 2012; Lindgaard et al., 2006. Exposure conditions measure first impressions, not task success.
This creates an asymmetry.
Beauty can receive evidence from color, contrast, proportion, rhythm, imagery, familiarity, and craft almost at once. Usability needs a person, a goal, an interface, a context, and enough time for an outcome.
ISO 9241-11 defines usability through effectiveness, efficiency, and satisfaction for specified users who pursue specified goals in specified contexts of use. The definition is not glamorous. It is useful because it blocks a common category error: a visual judgment and a quality-in-use judgment do not observe the same thing.
Beauty and usability are not enemies
The correction is not to make websites plain, joyless, or aggressively “functional.” Beauty is functional in several ways.
Rolf Reber, Norbert Schwarz, and Piotr Winkielman connect aesthetic pleasure with processing fluency. Symmetry, contrast, familiarity, repetition, and prototypicality can make a stimulus easier to process. A coherent visual system can reduce perceptual work. It can make grouping clearer, hierarchy faster to scan, and quality easier to trust.
Beauty is also more complex than “looks good.” Moshagen and Thielsch developed a website-aesthetics measure across seven studies and separated four facets: simplicity, diversity, colorfulness, and craftsmanship. Seckler, Opwis, and Tuch ran five online experiments and found that structural factors affected several aesthetic facets more broadly than color factors.
The useful conclusion is not that appearance is superficial. It is that appearance has its own mechanisms and measures. Those mechanisms can support use, but they do not contain all of use.
A beautifully grouped form can make relationships clear. It cannot make an ambiguous eligibility rule clear. A strong color hierarchy can make the primary action visible. It cannot tell the user what the action will do. Fine typography can invite reading. It cannot restore text that disappears at 200 percent zoom.
Beauty is part of the system. It is not a waiver for the rest of the system.
The aesthetic halo is real, but it is not armor
The phrase “what is beautiful is usable” came from serious research. It became misleading when people removed the conditions.
Noam Tractinsky, Adi Katz, and Dror Ikar used an ATM surrogate and found a strong relationship between perceived beauty and perceived usability before and after use. This matters. People do not put visual appeal in one sealed mental drawer and interaction quality in another.
But the relationship is not a universal law.
Tuch and colleagues assigned 80 participants to four online-shop conditions that independently varied aesthetics and usability. Beauty shaped pre-use impressions. After use, poor usability reduced perceived beauty. The interface did not merely fail while remaining visually untouched in the user's mind. Friction revised the visual judgment.
Sonderegger and Sauer studied 60 adolescents using two functionally identical simulated phones with different visual treatments. The appealing treatment improved perceived usability and reduced task-completion time in that context.
Beauty can create a halo, improve performance, or lose its shine
Actual use changes the relationship, and the result depends on what people do, what researchers measure, and where the interface fails.Tractinsky et al.
- Before use
- Beauty raised expected usability
- After use
- The perceptual halo remained after use
Tuch et al.
- Before use
- Aesthetics and usability were varied independently
- After use
- Poor use reduced later beauty ratings
Sonderegger & Sauer
- Before use
- Visual treatment changed perceived ease
- After use
- Appeal also shortened task time in this context
Different designs and tasks produced different relationships. None supports “beauty always equals usability.”
Tractinsky et al., 2000; Tuch et al., 2012; Sonderegger and Sauer, 2010. Results are intentionally not pooled.
These findings can coexist.
Beauty can create a halo. It can reduce anxiety, improve fluency, and sometimes improve performance. Strong friction can also puncture the halo. The direction depends on the interface, task, audience, measure, and moment.
That is why a design review that ends with “this feels intuitive” proves very little. The feeling may be genuine. The task evidence has not arrived.
Users predict the page before they operate it
Before a person clicks, they make a prediction.
Where is search? Which object is navigation? Does the logo return home? Is “Plans” a pricing page or a project-planning tool? Does this arrow open a panel, submit a form, or move the page?
These predictions rely on prior experience and local cues.
Roth and colleagues ran a preliminary study with 136 participants and a main study with 516 participants. They found robust, page-type-specific expectations for the locations of common web objects. Miniukovich and Figl studied 1,530 participants and more than 3,000 webpages. Prototypicality affected aesthetics, pre-use usability, and trustworthiness, and the relationship varied by website category.
A novel layout charges the user before the first click
People bring expectations about where common objects belong, so visual originality must repay the search and interpretation cost it creates.top-left
top-center
top-right
center
lower-right
Two studies mapped robust, page-type-specific expectations for common web objects.
Expected location is not a permanent law. It is a prediction cost that novel layouts must repay.
Roth et al., 2010; Miniukovich and Figl, 2023; Chi et al., 2001. Positions are representative rather than universal.
Conventions are not commandments. A creative site can move search, replace a navigation pattern, or invent a new interaction. But novelty creates a debt. The design must repay that debt with stronger cues, clearer consequences, or a benefit that makes the extra learning worthwhile.
The same rule applies to language.
Chi, Pirolli, Chen, and Pitkow modeled web navigation through information scent: cues help people estimate whether a path is likely to lead toward the information they need. A link can be beautifully styled and still have weak scent. “Explore,” “Discover,” and “Solutions” may fit the brand voice while forcing the user to guess.
When the guess is cheap, a person may continue. When the task involves money, health, risk, or a deadline, weak scent is not mysterious elegance. It is friction.
Friction accumulates after the first click
“Can people find the button?” is one usability question. It is not the whole interaction.
A useful diagnosis follows at least six moments:
- Predict. Can the person estimate where the relevant path begins?
- Orient. Do the available cues distinguish that path from nearby alternatives?
- Choose. Can the person predict what the control will do?
- Act. Can they operate it with their device, input mode, and ability?
- Confirm. Does the system show that the action succeeded, failed, or remains in progress?
- Recover. Can the person reverse, correct, or continue after a problem?
The first click is only one of six chances to get lost
A polished surface may reduce initial perceptual effort while labels, operation, feedback, and recovery make the task progressively harder.Predict
Where should the action be?
Orient
Which cue matches the goal?
Choose
What will this control do?
Act
Can the control be operated?
Confirm
Did the system register it?
Recover
Can the error be reversed?
Friction compounds. A polished surface can lower the first cost while labels, feedback, access, and recovery raise the next five.
Editorial synthesis from ISO 9241-11, information-scent research, usability measurement, and WCAG 2.2.
Visual fluency can reduce cost in the first three moments. It can help later too. Yet each moment has its own failure modes.
A tidy navigation can use labels that sound alike. A clear button can have a target that is too small. A smooth transition can hide a loading state. A minimal form can remove the instructions needed to correct an error. A confirmation can be visually elegant and semantically vague.
The costs compound because the user must carry uncertainty forward. If they are unsure whether the first action worked, the second action becomes a test of the interface rather than progress toward the goal. If recovery removes entered data, the person pays for the same information twice.
Tuch, Bargas-Avila, Opwis, and Wilhelm connected visual complexity with cognitive, emotional, physiological, performance, and memory measures. Their work is a reminder that visual structure matters. It is also a reminder that experience has several outputs. One screenshot cannot show them all.
Accessibility can fail inside a harmonious composition
A page can appear balanced and remain unusable for people who do not experience it under the designer's conditions.
WCAG 2.2 organizes accessibility requirements under four principles: perceivable, operable, understandable, and robust. It contains 13 guidelines and adds nine success criteria in version 2.2. The additions include focus visibility, target size, alternatives to dragging, consistent help, redundant entry, and accessible authentication.
These are not aesthetic preferences.
A pale interface may preserve a refined palette while failing contrast. A custom menu may fit the composition while trapping keyboard focus. A drag interaction may feel direct while excluding someone who cannot perform the gesture. A narrow column may look editorial at one viewport and become a horizontal-reading exercise under reflow.
Accessibility testing does not replace usability testing. It prevents the team from calling an experience usable when the task is structurally unavailable to part of the audience.
That distinction also explains why “our users did not complain” is weak evidence. People abandon, improvise, ask someone else, switch devices, or decide the service is not for them. Silence is not successful completion.
Diagnose the interaction, not the screenshot
The correct response to a beautiful but difficult site is not an immediate redesign. First, locate the failure.
Begin with a small set of representative tasks. Each task should name:
- the user group;
- the starting condition;
- the goal;
- the context and device;
- the completion condition; and
- the important failure conditions.
Then combine four evidence modes.
Technical and accessibility checks
Use automated and manual checks to identify contrast, semantics, keyboard access, focus, target, reflow, form, and performance failures. Automation is efficient where the condition is machine-testable. It cannot decide whether “Workspace” is meaningful to the customer.
Expert inspection
Nielsen and Molich tested heuristic evaluation in four experiments. Individual evaluators found only 20 to 51 percent of the identified usability problems, while aggregates of three to five evaluators performed well. One expert can find useful issues. One expert should not become the user population.
Representative task observation
Watch people attempt the task without teaching the interface. Record completion, critical errors, time or effort, hesitation, backtracking, abandonment, and recovery. Ask what they expect before the click and what they think happened after it.
Post-task perception
Measure confidence, perceived ease, satisfaction, and aesthetic response after the task. Pre-use and post-use ratings answer different questions. The difference between them can reveal where the visual promise met or lost the interaction.
A screenshot review is one thin layer of the evidence stack
Reliable diagnosis combines multiple evaluators, representative tasks, interaction outcomes, technical checks, and accessibility criteria.Problems found by one evaluator
Nielsen & Molich, 1990Evaluators in a stronger aggregate
Nielsen & Molich, 1990Usability studies reviewed
Hornbæk, 2006Reviewed studies without an interaction outcome
Hornbæk, 2006WCAG principles / guidelines / new 2.2 criteria
W3C, 2024Use several evidence modes: technical checks, expert inspection, representative tasks, and post-task perception.
Nielsen and Molich, 1990; Hornbæk, 2006; W3C, 2024. Values have different units and are not combined.
Kasper Hornbæk reviewed 180 HCI studies. Roughly one quarter did not assess the outcome of user interaction. Even formal usability work can measure partial indicators and miss quality in use.
The business version of the same mistake is common. A team measures visual preference, page speed, conversion, or satisfaction and calls the result “UX.” Each measure can be valuable. None, alone, explains whether the intended audience completed the intended task, at an acceptable cost, under the conditions that matter.
The best interface earns beauty twice
The first kind of beauty is immediate. It is proportion, type, color, motion, imagery, rhythm, and craft.
The second kind arrives through use. The label means what the person expected. The action works. The feedback removes doubt. The interface survives zoom, keyboard, interruption, error, and return. The system respects the user's time.
The two forms can reinforce each other. A strong visual system can make the route legible. A strong interaction can make the composition feel more considered after the task is done.
That is the more useful reading of the aesthetic-usability effect. Beauty is not decoration pasted onto use, and usability is not a penalty imposed on creativity. Both are judgments produced by the same person at different moments with different evidence.
The goal is not to choose one.
The goal is to ensure that the promise made in the first 50 milliseconds survives the next five minutes.
References
- Chi, E. H., Pirolli, P., Chen, K., & Pitkow, J. (2001). Using information scent to model user information needs and actions on the Web.
- Hornbæk, K. (2006). Current practice in measuring usability: Challenges to usability studies and research.
- ISO 9241-11:2018. Usability: Definitions and concepts.
- Lindgaard, G., Fernandes, G., Dudek, C., & Brown, J. (2006). Attention web designers: You have 50 milliseconds to make a good first impression!
- Miniukovich, A., & Figl, K. (2023). The effect of prototypicality on webpage aesthetics, usability, and trustworthiness.
- Moshagen, M., & Thielsch, M. T. (2010). Facets of visual aesthetics.
- Nielsen, J., & Molich, R. (1990). Heuristic evaluation of user interfaces.
- Reber, R., Schwarz, N., & Winkielman, P. (2004). Processing fluency and aesthetic pleasure.
- Roth, S. P., Schmutz, P., Pauwels, S. L., Bargas-Avila, J. A., & Opwis, K. (2010). Mental models for web objects.
- Seckler, M., Opwis, K., & Tuch, A. N. (2015). Linking objective design factors with subjective aesthetics.
- Sonderegger, A., & Sauer, J. (2010). The influence of design aesthetics in usability testing.
- Tractinsky, N., Katz, A. S., & Ikar, D. (2000). What is beautiful is usable.
- Tuch, A. N., Bargas-Avila, J. A., Opwis, K., & Wilhelm, F. H. (2009). Visual complexity of websites.
- Tuch, A. N., Presslaber, E. E., Stöcklin, M., Opwis, K., & Bargas-Avila, J. A. (2012). The role of visual complexity and prototypicality regarding first impression of websites.
- Tuch, A. N., Roth, S. P., Hornbæk, K., Opwis, K., & Bargas-Avila, J. A. (2012). Is beautiful really usable?
- W3C. (2024). Web Content Accessibility Guidelines (WCAG) 2.2.
Summary
Treat beauty as one useful part of the experience, then test whether representative people can predict, complete, confirm, and recover from the tasks that matter.
- Name the users, goals, contexts, and successful outcomes before reviewing the interface.
- Observe whether people can predict where key objects and actions should be.
- Check labels and local cues for strong information scent before changing the visual treatment.
- Measure completion, errors, effort, confidence, and recovery instead of collecting preference alone.
- Test keyboard, focus, target, reflow, contrast, and comprehension barriers against WCAG 2.2.
- Use several evaluators and representative task sessions; do not rely on one screenshot review.
- Retest aesthetic perception after use because interaction friction can change how the same page looks.