O1 has disappointed on every test so far. O1mini is interesting.
Here's three tests (code generation, chemistry first-principles solving and diagram generation).
Next, Remember this?
O1 and O1mini both failed to generate correct mermaid diagrams. O1mini corrected w multiple rounds to make something that competes with Sonnet, O1 failed multiple rounds of reflection (took about a minute per round).
x.com/hrishioa/status/1816331487182295241
Next is a clean @nextjs project with @shadcn components installed, a small JSON file (of my commits to repos) and asking for a dashboard.
Sonnet worked instantly. O1 & o1mini both needed multiple rounds. Eventually the o1 output was comparable to sonnet (sonnet first)
Maybe I'm missing something here, but is this just a model with additional CoT tokens (and some training on brain teasers) but no step-up in intelligence?
Will run more tests later.