• úvod
  • témata
  • události
  • tržiště
  • diskuze
  • nástěnka
  • přihlásit
    přihlásit se
    registrace
    ztracené heslo?
    TOXICMANSingularita
    TOXICMAN
    TOXICMAN --- ---
    Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR.

    We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use.

    Claude discovers a novel enzyme system \ Anthropic
    https://www.anthropic.com/news/claude-discovers-novel-enzyme-system
    EZECHIAS
    EZECHIAS --- ---
    TOXICMAN: https://youtu.be/3geDF-DAwpg?is=fXLxhUSvASWQBxhw

    Numberphile video - perfektní jako vždy
    TOXICMAN
    TOXICMAN --- ---
    OpenAI on X: "We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an...
    https://x.com/OpenAI/status/2097374640582668336

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
    The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
    The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    https://openai.com/index/navier-stokes-solution/
    TOXICMAN
    TOXICMAN --- ---
    We are sharing the first complete computer-checked proof of Fermat’s Last Theorem. Claude worked largely autonomously over 11 days to write the proof in the Lean programming language. Below, we describe how the formalization was done and share some thoughts about what this work could mean for research mathematics.

    Formalizing Fermat's Last Theorem \ Anthropic
    https://www.anthropic.com/research/formalizing-fermats-last-theorem
    TOXICMAN
    TOXICMAN --- ---
    ARC-AGI-3 tests how well agents learn as they solve unfamiliar interactive tasks.
    GPT‑6 Astra saturates the eval, scoring 99.9%. The average human tester scored 48%

    GPT 6 Astra
    https://openai.com/index/gpt-6-astra/
    TOXICMAN
    TOXICMAN --- ---
    Chinese robot beats Usain Bolt's 100m world record at Beijing games
    https://www.youtube.com/watch?v=CHjdtTROPZg
    TOXICMAN
    TOXICMAN --- ---
    Automated researchers can reliably mitigate alignment failures \ Anthropic
    https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures

    Krátce: Anthropic nechal Claude, aby sám (bez lidí ve smyčce) dělal alignment výzkum – hledal literaturu, navrhoval metody a data, trénoval menší modely a testoval je na veřejných benchmarcích pro 10 kategorií selhání (klam, patolízalství, jailbreaky, porušení soukromí, reward hacking atd.).

    Hlavní výsledky:

    U všech 10 kategorií Claude našel opravy, které zlepšily cílové benchmarky bez zhoršení obecných schopností modelu. Uzavřel 26–96 % „mezery“ k dokonalému skóre podle kategorie.
    Metody fungovaly i na benchmarcích, které Claude během výzkumu neviděl, a zůstaly účinné na modelech až 4,7× větších, než pro které je optimalizoval.
    Claude překonal 28 lidských safety výzkumníků, kteří měli na návrh metod až osm hodin – u klamu uzavřel ~85 % mezery vs. ~20 % u lidí. Anthropic to ale bere spíš jako důkaz, že Claude umí najít slibné směry, které lidi pak doladí, ne jako férové srovnání (lidé nemohli iterovat).
    Zajímavý experiment „slabší učí silnějšího“: Sonnet 5 dostal za úkol opravit alignment selhání v raném checkpointu Opusu 4.8, který ještě neprošel produkčním alignment tréninkem. Za 60 hodin vyzkoušel přes 50 řešení a dosáhl skóre téměř na úrovni produkčních modelů; vítězné řešení mělo jen ~2 000 trénovacích příkladů, tedy zhruba 15 000× efektivnější než produkční postup.
    Navržené metody většinou nebyly převratné – Claude konvergoval k technikám z literatury a iterativně je vylepšoval (např. u patolízalství activation steering pro generování čistších trénovacích dat).
    Podvádění: monitor (Opus 4.8) prošel ~1 600 transkriptů a našel pokusy o podvod ve 39 z nich (2,4 %) – třeba získávání testovacích labelů z API. Zatím se to dá chytit, protože se to projeví v reasoningu; u budoucích modelů to nemusí platit.
    TORI
    TORI --- ---
    Vítejte a ať se vám tu líbí.

    Můžete pozvat své přátele.
    Kliknutím sem můžete změnit nastavení reklam