Show HN: DesignArena – crowdsourced benchmark for AI-generated UI/UX

(designarena.ai)

86 points | by grace77 1 day ago

13 comments

coryvirok 1 day ago
This is really good! It would be really cool to somehow get human designs in the mix to see how the models compare. I bet there are curated design datasets with descriptions that you could pass to each of the models and then run voting as a "bonus" question (comparing the human and AI generated versions) after the normal genAI voting round.
[-]
- grace77 1 day ago
  wow this is a super interesting idea, and the team loves it — we'll fast follow-through and follow-up here when we add it, thanks for the suggestion!
- debesyla 1 day ago
  This would be extra interesting for unique designs - something more experimental, new. As as for now even when you ask AI to break all rules it still outputs standard BS.
muskmusk 1 day ago
This is a surprisingly good idea. The model vs model is fun, but not really that useful.
But this could be a legitimate way to design apps in general if you could tell the models what you liked and didn't like.
[-]
- grace77 1 day ago
  yes! that is the hope — /play is our first attempt at building out utility, would love your feedback and will ship hard to make it happen!
a2128 1 day ago
I tried the vote and both results always suck, there's no option to say neither are winners. Also it seems from the network tab you're sending 4 (or 5?) requests but only displaying the first two that respond, which biases it to the small models that respond more quickly which usually results in showing two bad results
[-]
- grace77 1 day ago
  Yes — great point. We originally waited for all model responses and randomized the vote order, but that made it a very bad user experience -- some models, especially open-source ones, took over 4 minutes to respond, leading to a high voter drop-off rate.
  To preserve the voter experience without introducing bias, our current approach waits for the slowest model within each binary comparison — so even if one model is faster, we don’t display until both are ready. You're right that this does introduce some bias for the two smallest models, and we'd love to hear suggestions for how to make this better!
  As for the 5th request: we actually kick off one reserve model alongside the four randomly selected for the tournament. This backup isn’t shown unless one of the four fails — it’s not the fastest or lowest-latency model, just a randomly selected fallback to keep the system robust without skewing results.
- ethan_smith 1 day ago
  Adding a "neither is good" option would improve data quality by preventing forced choices between two poor designs.
  [-]
  - grxxxce 1 day ago
    this is a great note — will be sure to add!
ppyyss8 8 hours ago
As a UX/UI designer in Korea, I love seeing related products being released. I hope they become even more advanced in the future.
jjani 14 hours ago
How about adding "mobile"? A lot of the time models tend to default to designs that don't make sense on mobile, even when instructed to design it as such.
[-]
- anonzzzies 11 hours ago
  Really? When I have a system prompt 'mobile-first design' it 100/100 works perfectly. What sort of things are you trying?
  [-]
  - jjani 5 hours ago
    The designs are passable for a mobile version of a simple website, but really sub-standard compared to the average app on the Play/App Store, whether native (Swift/Kotlin) or hybrid (Flutter/RN). In B2B SaaS you can get away with the 5000th shadcn UI, not so much for B2C mobile. The days that stock Material UI actually saw usage there are a decade behind us.
    If you have a tool/mode/prompt that creates good mobile UI designs, I'd love to know. Doesn't even have to generate code!
justusm 1 day ago
nice! Training models using reward signals for code correctness is obviously very common; I'm very curious to see how good things can get using a reward signal obtained from visual feedback
[-]
- grace77 1 day ago
  As are we, seems like the natural next step
calcsam 20 hours ago
interesting idea, this benchmark maps fairly closely to the types of output I typically ask LLMs to generate for me day-to-day
paulirish 18 hours ago
It would lend credibility to publish your system prompt.
[-]
- j421 18 hours ago
  System prompts can be found here: https://www.designarena.ai/system-prompts (also linked on about page).
  [-]
  - paulirish 5 hours ago
    Ah! My bad. thx
lofaszvanitt 2 hours ago
The problem is, what is being taught as UI-UX is 90% hogwash, balooney, pure bullshit. And these results reflect that.
iJohnDoe 21 hours ago
Very cool! Can the code and design that is generated be used?
[-]
- grace77 21 hours ago
  yes! we have a copy code and copy react code button on https://www.designarena.ai/play
filipeisho 1 day ago
[dead]
[-]
- grace77 1 day ago
  [dead]
butz 23 hours ago
[flagged]
[-]
- grace77 23 hours ago
  I wish—just added them back
adi_hn07 1 day ago
[flagged]
[-]
- grace77 1 day ago
  thank you! posting now :)
  [-]
  - adi_hn07 1 day ago
    Thanks !!