Typefully

Evaluation of O1 and O1mini: Tests and Comparisons

Avatar

Share

 • 

2 years ago

 • 

View on X

O1 has disappointed on every test so far. O1mini is interesting. Here's three tests (code generation, chemistry first-principles solving and diagram generation).
First we have a question from @Water_Splitting. Sonnet comes pretty close (I'll leave it to him for the details), O1 is WAY OFF. chatgpt.com/share/66e3e8b8-5278-8008-a073-16337b1e5643
Next, Remember this? O1 and O1mini both failed to generate correct mermaid diagrams. O1mini corrected w multiple rounds to make something that competes with Sonnet, O1 failed multiple rounds of reflection (took about a minute per round). x.com/hrishioa/status/1816331487182295241
Next is a clean @nextjs project with @shadcn components installed, a small JSON file (of my commits to repos) and asking for a dashboard. Sonnet worked instantly. O1 & o1mini both needed multiple rounds. Eventually the o1 output was comparable to sonnet (sonnet first)
Maybe I'm missing something here, but is this just a model with additional CoT tokens (and some training on brain teasers) but no step-up in intelligence? Will run more tests later.
Avatar

Hrishi Olickel

@hrishioa

Building artificially intelligent bridges at Southbridge, prev-CTO Greywing (YC W21). Chop wood carry water.