reading

things i've saved from my curius. a curated subset lives in the bookshelf.

  • Horses

    andyljones.com

    AI progress is steady. Human equivalence is sudden.

  • theatlantic.com

    Why reactionaries are taking over the world

  • experimental-history.com

    Why do we have to learn everything the hard way?

  • masonjwang.com

    My notes from Stanford’s incredible course on World War 2.

  • sive.rs

    Derek Sivers official site. Thoughts on philosophy, culture, self-improvement. Author of Useful Not True, How to Live, Hell Yeah or No, Anything You Want.

  • archive.nytimes.com

    Does Joseph Smith’s theology suggest that one of our presidential candidates could be a deity?

  • scottaaronson.blog

    As most readers have presumably heard by now, Paul Erdös’s Unit Distance Problem from 1946—one of the central open problems from the field of discrete geometry—has been solved by …

  • thefp.com

    The work of building frontier AI has brought us to the edge of where He might be, writes Avital Balwit.

  • youtube.com

    Dwarkesh Clips explores the architectural differences between GPUs and TPUs by examining their distinct high-level block structures. The discussion focuses on how each design manages data movement, core organization, and hardware efficiency to support complex computational tasks.

  • sriramk.com

    Any large system picks a metric to goal itself on. Entire books and way-too-long Medium posts have been written on the importance of said metric - it influences everything from people’s incentives to how quickly you can optimize your business. In an organizational equivalent of Schrödinger’s cat, picking the metric itself can cause weird cultural distortion (see Goodhart’s Law). Since it is near impossible to perfectly measure human behavior, most large teams/products pick a proxy metric to measure underlying behavior. For example - ‘clicks’ are a proxy for “did I read this?” and “will I buy this product sometime in the future?”, ‘time spent’ is a proxy for “did I enjoy this content?” and NPS is often a substitute for “do I love this company?”. You convert a nebulous human emotion/behavior to a quantifiable metric you can align execution on and stick on a graph and measure teams on. Engineers and data scientists can’t do anything with “this makes people feel warm and fuzzy”. They can d

  • sriramk.com

    I’ve been discussing with some frontier lab researcher friends as to why frontier model capabilities are always so clustered together as opposed to any one model having an unassailable edge. The best metaphor for this in my mind is the “four minute mile”: no one broke it till Bannister in 1954 and then very quickly five more runners did it in the next two years. In the model world, this translates to: a) Intense competitive pressure. b) Often similar pool of ideas and research directions. c) Roughly similar access to capitalization and compute infrastructure. The oft quoted “fast follow” example is after the launch of o1 being quickly followed by reasoning models from multiple players, both closed and open weights. This is not the case with many other technology driven industries where capability or advancements often tend to be longer held and a fast-follow model is harder.

  • evjang.com

    A TUTORIAL Automating Go Research with AutoGo Building a strong Go AI from scratch with modern AI tools. PLAY AT AUTOGO.EVJANG.COM → CODE ON GITHUB → AUTOGO — cover — how to navigate 01 · motivation 02 · thanks INTRO TO GO ALPHAGO TUTORIAL RESEARCH FINDINGS AUTOGO · ERIC JANG

  • personalsit.es

    Personal sites are sick as hell, so this site was built so we can all discover each other's. This directory of links are by folks that want to share their site with the world.

  • thoughts.melonking.net

    There’s a simple test you can do to tell if a website is a positive citizen of the web, or a negative one. Go to the website and look for their links; do they have any? Does the website link to other websites made by other people? Or do they just link to their own social media? How many outbound links do they have?

  • henry.codes

    On the eve of the new year, 2025, I was possessed by the spirit of adventure, and drove to Wyoming in the middle of the night. Hijinks, as is eternally their way, ensued.

  • aman.ai

    Course notes and learning material for Artificial Intelligence and Deep Learning Stanford classes.

  • THE METIS LIST

    metislist.com

    THE WORLD'S TOP AI RESEARCHERS CREATED WITH FIGMA MAKE LAST UPDATED: JULY 28 9:04PM SSI CITATIONS: 663,198 UNDERGRAD: UNIVERSITY OF TORONTO PHD: UNIVERSITY OF TORONTO PREVIOUS COMPANY: OPENAI DWARKESH APPEARANCE: YES NEURIPS/ICLR/ICML: YES INTERESTS AND NOTABLE WORKS: ALEXNET, SEQ2SEQ, DEEP LEARNING GOOGLE DEEPMIND CITATIONS: 262,612 UNDERGRAD: DUKE UNIVERSITY PHD: NONE PREVIOUS COMPANY: CHARACTER AI DWARKESH APPEARANCE: YES NEURIPS/ICLR/ICML: YES INTERESTS AND NOTABLE WORKS: ATTENTION IS ALL YOU NEED, MOE, CHARACTER.AI UNIVERSITY OF TORONTO CITATIONS: 937,866 UNDERGRAD: UNIVERSITY OF CAMBRIDGE PHD: UNIVERSITY OF EDINBURGH PREVIOUS COMPANY: GOOGLE DWARKESH APPEARANCE: NO NEURIPS/ICLR/ICML: YES INTERESTS AND NOTABLE WORKS: BACKPROP, BOLTZMANN MACHINE, DEEP LEARNING THINKING MACHINES UNDERGRAD: OLIN COLLEGE PHD: NONE PREVIOUS COMPANY: OPENAI NEURIPS/ICLR/ICML: YES CITATIONS: 254,381 DWARKESH APPEARANCE: NO INTERESTS AND NOTABLE WORKS: GANS, GPT, CLIP MIT UNDERGRAD: TSINGHUA UNIVERSITY PH

  • interconnects.ai

    The cutting edge of AI, from inside the frontier AI labs, minus the hype. The border between high-level and technical thinking. Read by leading engineers, researchers, and investors. Click to read Interconnects AI, a Substack publication with tens of thousands of subscribers.

  • Yi Tay

    yitay.net

    Documenting my 3.3 years at Google Research and Brain. Releasing the new open source Flan-UL2 20B model. Here are some of the best language AI / NLP papers of 2022! Some thoughts on emergent abilities and scaling language models. Made with Squarespace

  • founders.school

    Forty weeks building the kind of person who can run a million-dollar company. Elite entrepreneurs who think deeply, build with AI, and run real businesses.

  • vladfeinberg.com

    Vlad's Blog

  • arxiv.org

    Abstract:The push to train ever larger neural networks has motivated the study of initialization and training at large network width. A key challenge is to scale training so that a network's internal representations evolve nontrivially at all widths, a process known as feature learning. Here, we show that feature learning is achieved by scaling the spectral norm of weight matrices and their updates like $\sqrt{\texttt{fan-out}/\texttt{fan-in}}$, in contrast to widely used but heuristic scalings based on Frobenius norm and entry size. Our spectral scaling analysis also leads to an elementary derivation of \emph{maximal update parametrization}. All in all, we aim to provide the reader with a solid conceptual understanding of feature learning in neural networks.

  • acotra.substack.com

    The kids are alright

  • neuroscience.stanford.edu

    Join the speaker for coffee, cookies, and conversation before the talk, starting at 11:45am.Beyond the

  • acotra.substack.com

    Weigh the costs against the benefits

  • noahpinion.blog

    All warfare is drone warfare now.

  • sfbestof.com

    A curated map of my favorite places in San Francisco.

  • github.com

    A curated list of best cuda programming books

  • near.blog

    Complex systems of life often contain a multitude of shortcuts: sources of alpha which, should you choose to exploit them, give notable advantage. Many classes of these shortcuts are features rather than bugs, and it is no accident of many systems that only a select minority are able to tactfully navigate them. In many cases there exists zero-sum shortcuts, which, if all of society were to adopt them tomorrow, would have their effectiveness instantly curtailed to zero. Luckily there also exists many which are positive-sum and benefit both parties involved. This post contains a few notes and examples from both classes. In venture capital, it is strongly preferred that one receives an introduction to an investor prior to pitching them. This is, in general, a much easier way to meet with many classes of professionals than a cold-email. At first glance this may appear shallow or nepotistic, but it’s more accurate to view it as one of the first tests to becoming a successful founder. If you

  • near.blog

    I find the above examples fascinating from the meta perspective: while there’s nothing wrong with having fun building inside of games, these are the very same skillsets which tech companies would pay six figures to have on their side! (of course they may enjoy the games more – this is discussed later) Sometimes people of this caliber even have trouble finding a job – they often don’t really know where to go besides apply online to boomer companies who reject them when they see a lack of credentials. They love building things and are very smart and hardworking, but their milieu is an environment which captured them from a young age (often a video game or social media) and sometimes also ensured that the value they produce is within a pre-existing platform (e.g. a video game). “Why did you cherry-pick people playing video games instead of talking about, like, everyone currently enrolled in medical school or something?” The US currently has 125,000 students enrolled in medical school (

showing 30 of 396