<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://skumar-ml.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://skumar-ml.github.io/" rel="alternate" type="text/html" /><updated>2026-09-09T22:09:57+00:00</updated><id>https://skumar-ml.github.io/feed.xml</id><title type="html">Home</title><subtitle>Personal website of Your Name, PhD Student in Your Field at Your University.</subtitle><entry><title type="html">PhD Students Need Financial Advice Too</title><link href="https://skumar-ml.github.io/blog/financial-advice-for-grad-students/" rel="alternate" type="text/html" title="PhD Students Need Financial Advice Too" /><published>2026-07-14T17:00:00+00:00</published><updated>2026-07-14T17:00:00+00:00</updated><id>https://skumar-ml.github.io/blog/financial-advice-for-grad-students</id><content type="html" xml:base="https://skumar-ml.github.io/blog/financial-advice-for-grad-students/"><![CDATA[<p>I was recently chatting with one of my intern-buddies, and I was surprised to learn he kept all of his cash in his bank account. Not a single penny was invested. I explained to him why he should invest, and after he came around to my view, I started to explain the basics of investing. That conversation made me remember how clueless I was when I first started investing in 2020, and it helped me realize how much implicit investing knowledge I had soaked up over the years.</p>

<p>But you’re a PhD student, and PhD students are busy! We have a bunch of things constantly pulling at our attention. You need to grade homework for the class you TA, your advisor wants your help writing a proposal, NeurIPS assigned you 5 papers to review, and the list never ends. If you’ve never invested before, it’s easy to let investing slip to the bottom of your priority list. It’s one of those things that feels daunting to start on because there are so many unknowns to navigate.</p>

<p>So, I’m writing this blog post to convince you to start investing, with some pointers on how to get started.</p>

<figure class="post-figure post-figure--half">
  <img src="/assets/images/OpenInvestingMeme.jpg" alt="Two-panel meme: Nicolas Cage looking stressed with text 'Some random guy asking me to start investing'; Pedro Pascal laughing with text 'Me who can barely survive on my stipend'." />
</figure>

<h2 id="1-why-phd-students-should-care-about-investing"><a href="#1-why-phd-students-should-care-about-investing"></a>1. Why PhD Students Should Care About Investing</h2>

<ol>
  <li>
    <p><strong>Investing during grad school is one of the lowest-risk ways to learn, arguably, the highest-leverage skill of your life</strong>. At some point, you will begin investing your money<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">I’m assuming you are convinced about the benefits of investing. If you do not believe in investing, arguments to convince you otherwise are <a href="https://awealthofcommonsense.com/2024/09/what-if-you-only-invested-at-market-peaks-2/" target="_blank" rel="noopener noreferrer">in</a> <a href="https://www.etf.com/docs/IfYouCan.pdf" target="_blank" rel="noopener noreferrer">these</a> <a href="https://jlcollinsnh.com/2012/04/19/stocks-part-ii-the-market-always-goes-up/" target="_blank" rel="noopener noreferrer">pieces</a>.</span>
</span>
. Right now, your yearly earnings are (hopefully) at the lowest point they will ever be in your career. That’s why now is the perfect time to invest. You <strong>will</strong> make mistakes, and there will be days when your investments are in the red (i.e., you are losing money). Even though it hurts to lose <span class="tex2jax_ignore">$100</span> when your PhD only pays you <span class="tex2jax_ignore">$2,800</span> a month, these mistakes are a <em>lot</em> less painful now than in the future, when your earnings have increased 5-20x. At that point, you want to know what you are doing, not figuring it out for the first time.</p>
  </li>
  <li>
    <p><strong>Experience is the greatest teacher</strong>. You can read all the Reddit posts and watch all of the YouTube videos you want. None of it will get you ready for the emotional rollercoaster of being invested in the market. There’s no way to tell what your emotional reaction will be during a market crash, when your portfolio is down 10-20%—whether you’ll be able to resist panic selling, or whether you’ll have the courage to invest more while the market is volatile<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Or if you are like me and continue to stay on the sidelines since you’re greedy enough to believe you can perfectly time the market.</span>
</span>
. Living these experiences firsthand will prepare you for when real money is on the line.</p>
  </li>
</ol>

<p>So, please, start investing. Your future you will thank you.</p>

<h2 id="2-what-to-invest-in"><a href="#2-what-to-invest-in"></a>2. What to Invest In</h2>

<p>There are a million options for what to buy. At the simplest level, you can invest in equities or bonds. Equities can either be <strong>stocks</strong> (individual publicly traded companies like Google, Tesla), or you can invest in an <strong>index fund</strong>, which is a weighted average of many companies. For example, the <a href="https://en.wikipedia.org/wiki/S%26P_500" target="_blank" rel="noopener noreferrer">S&amp;P 500</a> is a weighted average of the 500 largest companies in the USA. There are also index funds for specific categories (<a href="https://en.wikipedia.org/wiki/Nasdaq-100" target="_blank" rel="noopener noreferrer">NASDAQ-100</a> is predominantly for tech, <a href="https://investor.vanguard.com/investment-products/etfs/profile/vht" target="_blank" rel="noopener noreferrer">VHT</a> tracks the health care sector, etc).<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Technically, the S&amp;P 500 is an index, and there are various index funds (like VOO, IVV, SPYM) designed to closely track the original index.</span>
</span>
. A bond is a loan <em>you</em> give to the government, and in exchange they pay you interest. You typically need to hold a bond for a prespecified amount of time (can be from 1-30 years), and you typically forfeit some of the interest if you sell the bond.</p>

<p>If you believe America will generally do well over the coming decades, then index funds are a great choice. You average out the fluctuations of individual companies, which means you track the American economy (or the growth of a specific sector). If you feel like an individual company will win, you can opt to invest in its stock—this gives you a chance to beat index funds, but may expose you to more volatility. If you’re unsure, bonds are a safe bet (you are betting that the US government will be able to pay you back, which is one of the safest bets that can be made).</p>

<h2 id="3-investing-strategies"><a href="#3-investing-strategies"></a>3. Investing Strategies</h2>

<p>As a grad student, you realistically don’t have time to actively trade the market or do significant market research to pick individual winners. If you think you can outsmart Wall Street and find mispriced bets, you’re likely to be wrong<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">For more on this, check <a href="https://jlcollinsnh.com/2012/04/25/stocks-part-iii-most-people-lose-money-in-the-market/" target="_blank" rel="noopener noreferrer">this article</a>.</span>
</span>
. Institutional investors have access to resources that retail investors (jargon that refers to casuals like us) will never have, and this is literally their full-time job. But the benefit of living in a free market is that if you have a strong belief about a certain stock, you can go ahead and buy it<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Perhaps a keen AI researcher could have foreseen the explosion in NVIDIA’s stock.</span>
</span>
.</p>

<p>My personal preference is to follow a <strong>buy and hold strategy</strong>. In this strategy, once investments are made, they are rarely sold. The buyer believes that the investment will, on average, have good returns over a <em>long</em> period of time (on the order of years to decades). This is diametrically opposed to <strong>active trading</strong>, where one is buying and selling on a frequent basis (on the order of days to weeks).</p>

<p>A key concept to be aware of is <strong>Dollar-Cost Averaging (DCAing)</strong>. Let’s say you have <span class="tex2jax_ignore">$1,000</span> to invest, and you know where you want to invest it. Instead of investing it all today, you invest it over a span of time (e.g., invest <span class="tex2jax_ignore">$33</span> every day for a month, or <span class="tex2jax_ignore">$250</span> every week). The key idea is that the market has natural volatility, and so it’s natural for a clever investor to want to “time the market” (pick a point where the price of the investment is low). In practice, this is really hard to do. DCAing allows you to participate if the market rises, while also exposing you to entry points at lower prices if the market falls. The beauty is that you don’t even need to log in to your brokerage every day. Every modern brokerage has an automated DCA feature that you can set. It’s a great strategy. <a href="https://ofdollarsanddata.com/even-god-couldnt-beat-dollar-cost-averaging/" target="_blank" rel="noopener noreferrer">Even God can’t beat it</a>.</p>

<p>DCAing can work well for a lot of people. But some folks want more control, or think they can do better than DCAing (like myself). My personal strategy is to keep at most X% of my portfolio as cash in case there’s a market pullback, and DCA the rest<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">For me, X is around 30%.</span>
</span>
. This way, I stay exposed to gains in the market, while saving some dry powder in case there’s a more opportune moment to invest. There are a few downsides with this approach: (1) you actually have to track the market to watch for opportunities, and sometimes you can miss some if it’s a busy period in life; (2) when a pullback is actively happening, it’s easy to fall into the trap of trying to perfectly “time” the dip. Before you know it, the market can recover, and your opportunity is gone; you would have been better off just DCAing.</p>

<p>If I were a smarter man, I’d DCA all of my cash. But I’m foolish enough to believe I can do better.</p>

<figure class="post-figure post-figure--half">
  <img src="/assets/images/TimingMarketMeme.jpeg" alt="Distracted boyfriend meme: the man labeled Investors looks at Timing the Market while his girlfriend Buy and Hold looks on disapprovingly." />
</figure>

<h2 id="4-what-brokerage-to-invest-through"><a href="#4-what-brokerage-to-invest-through"></a>4. What Brokerage to Invest Through?</h2>

<p>You invest through a brokerage, which is a platform that provides investing services. <a href="https://robinhood.com/" target="_blank" rel="noopener noreferrer">Robinhood</a> is great for beginners. They have a nice UI, which makes investing feel much easier. Personally, I use <a href="https://www.fidelity.com/" target="_blank" rel="noopener noreferrer">Fidelity</a>. If you haven’t opened a brokerage account, ask your friends for a referral link. You’ll both get rewarded! Otherwise, you’re always welcome to use mine<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Robinhood referral link: <a href="https://join.robinhood.com/shubhak-9f74ea/" target="_blank" rel="noopener noreferrer">https://join.robinhood.com/shubhak-9f74ea/</a>. But seriously, check with your friends first.</span>
</span>
 (yes, this is a shameless plug).</p>

<p>If you’ve gotten to this point, you know the basics for investing. Feel free to stop reading and start investing! The rest of the blog post talks about more advanced knowledge I’ve accumulated over the years.</p>

<h2 id="5-taxes--the-roth-ira"><a href="#5-taxes--the-roth-ira"></a>5. Taxes &amp; the Roth IRA</h2>

<p>An individual brokerage account is the most basic way to invest (this is the default account you open with any brokerage). The money invested in this account has already been taxed (what we call <em>post-tax</em> money). If your investments grow (e.g., from 10k to 30k) and you sell those investments, you need to pay taxes on the earned 20k, which can be significant<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Federally, positions held for under a year are subject to income tax and positions held for over a year are generally taxed at a lower long-term capital gains rate (15%). At the state level, the rules vary; some states tax as income tax, others give long-term preferences, and some don’t tax at all.</span>
</span>
.</p>

<p>But what if there was a way to avoid paying tax? This is where <em>tax-advantaged</em> brokerage accounts come in. In particular, for grad students, I highly recommend a Roth IRA. When you’re a grad student, you are typically in the lowest tax bracket of your life. Thus, you want to pay taxes <em>now</em>, rather than later on when your salary puts you in a higher tax bracket<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Contrast this with when your earnings are at their peak. In such cases, you probably want to invest in a <em>pre-tax</em> IRA, which allows you to defer your taxes to later years, when your income may be less (for example, when you are retired).</span>
</span>
. Like with an individual account, you contribute post-tax money to a Roth IRA. When you sell investments or make <em>qualified</em> withdrawals, they are completely tax-free (no income tax or capital gains tax).</p>

<p>Yes, there’s a catch (actually, there are two). The amount you can contribute to a Roth IRA every year is capped (in 2026, it was <span class="tex2jax_ignore">$7,500</span>). Also, there are rules for withdrawing from the Roth IRA: (1) you can withdraw your principal (i.e., your contributions) at any time without penalty; (2) earnings can be withdrawn tax-free after you are 59.5 years old<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">There are a couple of other nuances that are worth being aware of. You can read more <a href="https://www.irs.gov/retirement-plans/roth-iras" target="_blank" rel="noopener noreferrer">here</a>.</span>
</span>
. If you are comfortable with parking your investments for the long term, then a Roth IRA can be one of the best ways to reduce your tax liability.</p>

<p>If you’re interested, your brokerage should have an option to open a Roth IRA. Personally, I try to max out my Roth IRA contributions every year, and then I invest any leftover money in my individual brokerage account.</p>

<h2 id="6-earning-interest-on-your-cash"><a href="#6-earning-interest-on-your-cash"></a>6. Earning Interest on Your Cash</h2>

<p>I try to keep as little cash as possible in my checking account. It’s recommended to keep enough to survive (rent + food) for 3-6 months and cover any small, surprise expense (e.g., unexpected travel or a night out)<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">If you have a strong safety net, then you can be more aggressive. For example, I keep around 2 months of expenses in my account, since I’m fortunate to have a financially stable family.</span>
</span>
. Then, I try to move as much cash as possible into a money market fund through my brokerage accounts (a money market fund essentially gives you interest on your cash, and you can pull the cash out at any time)<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Just be aware that if you ever need to access cash in your brokerage account, it can take 3-5 business days to transfer to your bank account. So make sure you keep enough cash in your checking account to cover your expected and unexpected short-term expenses.</span>
</span>
.</p>

<p>Interest rates vary over time (they are set by the Federal Reserve, which is like the bank for banks). When the interest rate is high, brokerages will give you a higher return on your cash. Even if it’s low, it beats having your cash sit in a checking account, where you’re probably getting zero interest. Currently, the interest rate is 3.5%. In recent years, it’s been as high as 5.5%.</p>

<h2 id="concluding-thoughts"><a href="#concluding-thoughts"></a>Concluding Thoughts</h2>

<p>Even if you decide not to follow any of the investing advice in this post (which is totally okay!), I hope I’ve convinced you to get started with investing (however you best see fit). Make your money work for you, so eventually, you don’t have to!</p>]]></content><author><name></name></author><category term="blog" /><summary type="html"><![CDATA[Investing as a PhD student is one of the lowest-risk ways to learn, arguably, the highest-leverage skill of your life. This post convinces you to start investing, with some pointers on how to get started.]]></summary></entry><entry><title type="html">We Need a Technical Definition for World Models</title><link href="https://skumar-ml.github.io/blog/a-technical-definition-of-world-models/" rel="alternate" type="text/html" title="We Need a Technical Definition for World Models" /><published>2026-06-27T17:00:00+00:00</published><updated>2026-06-27T17:00:00+00:00</updated><id>https://skumar-ml.github.io/blog/a-technical-definition-of-world-models</id><content type="html" xml:base="https://skumar-ml.github.io/blog/a-technical-definition-of-world-models/"><![CDATA[<div class="post-epigraph">
  <p class="post-epigraph-text">&ldquo;You see, but you do not observe. The distinction is clear.&rdquo;</p>
  <p class="post-epigraph-cite">&mdash; Sherlock Holmes to Dr. Watson, <em>A Scandal in Bohemia</em></p>
</div>

<p>Fads in AI come and go nearly as fast as (or arguably, faster than) fashion trends. CNNs, RNNs, transformers, diffusion, agents, and now world models. It feels like many papers I read lately slyly slip “world models” somewhere, and during my time at IBM this summer, I’ve had the good fortune of studying world models. <strong>My conclusion: most people are not actually building world models.</strong></p>

<p>This may be partly attributed to the loose definition of world models. It’s a loaded term that sounds flashy but currently lacks a <em>technical</em> definition (mainly because no one yet knows how to build one). All definitions for world models I have come across (both from classical thinkers and modern researchers) are high-level and conceptual<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">If you’ve come across other technical definitions, please let me know!</span>
</span>
. They help frame what a world model is and why they are helpful, but they do not explain what a world model technically does.</p>

<p><strong>This post seeks to technically define world models by listing the capabilities they should have (I believe there are three).</strong> Then, I examine popular approaches to world models (LeCun’s JEPA, Fei-Fei’s World Labs, Jim Fan’s World Action Models) and show that they are, according to my definition, not world models. I end with some forward-looking thoughts on where I think this all converges.</p>

<h2 id="so-what-is-a-world-model-anyway"><a href="#so-what-is-a-world-model-anyway"></a>So what is a world model anyway?</h2>

<p>A world model, by the most generic definition possible, is an “internal model of how the world works” (<a href="https://openreview.net/pdf?id=BZ5a1r-kVsf">LeCun</a>). It tells an agent what is (or what was/what would be) likely, plausible, and impossible (<a href="https://openreview.net/pdf?id=BZ5a1r-kVsf">LeCun</a>).</p>

<p>Let’s take a look at a real-life example. Recently I went to grab my lunch from a shared fridge. Upon opening the fridge, I realized that my lunch was not on the bottom row where I had originally placed it. What happened to it? My internal world model told me that either (A) someone stole it or (B) someone moved it. My world model judged (A) as unlikely, since I was in an IBM office (not some college dorm), but not impossible. To reduce the uncertainty, my internal <em>policy</em> directed me to gather more <em>observations</em> by looking elsewhere in the fridge, which led me to locate my lunch on the top shelf. Thus, my <em>belief</em> collapsed to (B). Note that my world model didn’t think of other explanations, such as (C) a deer entered the office, found the fridge, and ate my lunch. Strictly plausible, but highly unlikely<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">And if that had indeed occurred, my world model would probably be seriously updated.</span>
</span>
.</p>

<figure class="post-figure">
  <img src="/assets/images/lunch-world-model.png" alt="Three bar charts showing belief over lunch location: first certain on the fridge bottom, then uncertain between fridge top and gone, then certain on the fridge top." />
  <figcaption>How my belief over lunch location evolved using my world model: initially certain it was on the bottom shelf, uncertain when it was missing (probably somewhere else in the fridge, but maybe it was stolen), then certain again after I checked elsewhere in the fridge.</figcaption>
</figure>

<h2 id="a-technical-definition-of-world-models"><a href="#a-technical-definition-of-world-models"></a>A technical definition of world models</h2>

<p>Some terminology first. An <em>observation</em> is something that is seen by the agent. For embodied agents (e.g., robots and humans), this is sensory information (e.g., from eyes and ears). An <em>abstract state</em> is a higher-level, compact, and structured representation that is extracted from the observations. It is the useful information our world model operates on. For example, if one is driving, the abstract state could include location of pedestrians, other cars, speed limit, weather, etc. All of these are <em>inferrable</em> from our observations, but they are <em>not</em> the same as our observations.</p>

<p>Technically speaking, any system that meets the following criteria is a world model:</p>

<ul>
  <li><strong>Goal #1.</strong> Extract an abstract state from observations.</li>
  <li><strong>Goal #2.</strong> Update the abstract state based on new observations.</li>
  <li><strong>Goal #3.</strong> Predict the next abstract state, given the current state and an action.</li>
</ul>

<p>There are bonus desiderata that would make a world model more useful, but I don’t believe they should be part of the definition<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Having a semantic abstract state is beneficial, since one can interpret the abstract state. It also allows for policies to interface with a world model through natural language. Relatedly, having a reasoning interface can allow for question-answering about the abstract state, rather than using the world model only as a simulator.</span>
</span>
.</p>

<p>The world model is <em>not</em> responsible for:</p>
<ol>
  <li>Planning actions.</li>
  <li>Rendering a simulation of partial observations from the predicted next state.</li>
</ol>

<p>This directly contradicts <a href="https://drfeifei.substack.com/p/a-functional-taxonomy-of-world-models">Fei-Fei’s definition of a world model</a>. I think that what she calls a world model is really an amalgamation of a planner, renderer, and world model. At the end of the day, it’s possible that all three of these components get folded into one end-to-end model, but it is not correct to <em>solely</em> call such a model a world model (instead, such a model includes a world model).</p>

<h2 id="lecuns-jepa-based-world-models"><a href="#lecuns-jepa-based-world-models"></a>LeCun’s JEPA-based World Models</h2>
<p>LeCun has been building world models using the JEPA framework. Put simply, JEPA is a self-supervised way of training a backbone to accomplish Goal #1: given partial observations, a <em>state extractor</em> learns to represent the abstract state as a latent embedding. Then, an <em>action-conditioned next-state predictor</em> can be learned on top of the state extractor to accomplish Goal #3. Since LeCun’s components are modular, the state extractor and action-conditioned next-state predictor can be trained separately (see <a href="https://arxiv.org/pdf/2603.14482">V-JEPA</a>) or together in an end-to-end fashion (see <a href="https://arxiv.org/pdf/2603.19312">LeWorldModel</a>).</p>

<p>The major limitation of LeCun’s approach thus far is Goal #2. My personal opinion is that a more structured, latent <em>variable</em><span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Something that represents distributions over abstract state, rather than the point estimate of an embedding.</span>
</span>
 representation is needed to effectively reach this goal. From what I can tell, this seems to be something his group is actively working towards (see <a href="https://arxiv.org/pdf/2602.11389">C-JEPA</a>).</p>

<figure class="post-figure post-figure--triptych">
  <div class="post-figure-panels">
    <div class="post-figure-panel post-figure-panel--primary">
      <p class="post-figure-panel-label">LeCun (JEPA)</p>
      <img src="/assets/images/LeCun.png" alt="Diagram of LeCun's JEPA approach: an observation is encoded into an abstract state, which together with an action predicts the next abstract state." />
    </div>
    <div class="post-figure-panel-stack">
      <div class="post-figure-panel">
        <p class="post-figure-panel-label">Fei-Fei (World Labs)</p>
        <img src="/assets/images/FeiFei.png" alt="Diagram of Fei-Fei's World Labs approach: an observation is mapped to an interactive, editable 3D world." />
      </div>
      <div class="post-figure-panel">
        <p class="post-figure-panel-label">Jim Fan (WAM)</p>
        <img src="/assets/images/JimFan.png" alt="Diagram of Jim Fan's World Action Model: given an observation and past action, the model predicts the next action and next observation." />
      </div>
    </div>
  </div>
  <figcaption>A schematic comparison of three prominent &ldquo;world model&rdquo; approaches. LeCun&rsquo;s JEPA extracts abstract state and predicts the next abstract state; Fei-Fei&rsquo;s World Labs outputs an interactive 3D world; Jim Fan&rsquo;s WAM jointly predicts actions and next observations.</figcaption>
</figure>

<h2 id="fei-feis-world-models"><a href="#fei-feis-world-models"></a>Fei-Fei’s World Models</h2>
<p>It’s well known in the AI world that Fei-Fei (a legend for young computer vision researchers such as myself) has started <a href="https://www.worldlabs.ai/">World Labs</a>, and they are trying to build a world model for spatial intelligence. However, I don’t think they are actually building a world model.</p>

<p>In my view, they are really building a <em>3D simulation environment</em><span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Which is useful and non-trivial, but separate from a world model.</span>
</span>
. The 3D environment is essentially richer observed data, but it is <em>not</em> an abstract state<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">One could argue against this by redefining what abstract state means.</span>
</span>
. Given a living room, the abstract state would contain things like: the mug is dirty, the TV is on, the cat is on the couch. The 3D model certainly contains this information in its rendering, but one would still need an abstract state extractor. Perhaps World Labs’s argument is that accomplishing Goals #1–3 is much easier when you can learn a 3D simulator, because geometrical and physical constraints are enforced. Maybe that’s true, but it doesn’t mean they are building a world model.</p>

<p>Effectively, I’d argue that they are building a next-observable-state simulator, and on top of that representation, one would still need to accomplish Goals #1–3, which I don’t believe is trivial.</p>

<figure class="post-figure">
  <img src="/assets/images/world-labs-webp.webp" alt="Demo of World Labs generating an interactive, editable 3D environment." />
  <figcaption>A demo from the <a href="https://www.worldlabs.ai/" target="_blank" rel="noopener noreferrer">World Labs</a> homepage, showing a generated 3D world.</figcaption>
</figure>

<h2 id="jim-fans-wam"><a href="#jim-fans-wam"></a>Jim Fan’s WAM</h2>
<p>Jim Fan’s—and by extension NVIDIA’s—<a href="https://www.linkedin.com/pulse/second-pre-training-paradigm-jim-fan-xn5fc/">thesis</a> is that we can directly learn the policy and fold the world model into the learned policy. They are pushing World Action Models (WAMs). WAMs start with a pre-trained video-generation backbone, and train to predict future states (the state is represented by video frames) and robot actions when conditioned on past video frames, language instructions, and action sequences. Then, once the WAM is trained, it can run as a feedforward policy at inference time (the predicted image frames are typically discarded when using the WAM as a policy).</p>

<p>I think this approach is not world modeling for reasons similar to Fei-Fei’s approach. Video frames are <em>not</em> an abstract state. But Jim is probably betting that a policy doesn’t need an explicit world model<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">I would pay to watch LeCun and Jim duke this out.</span>
</span>
.</p>

<h2 id="so-where-does-that-leave-us"><a href="#so-where-does-that-leave-us"></a>So where does that leave us?</h2>

<p>By my definition, none of the “big three” are building world models. They are building models that are likely helpful for robotics, but I don’t see how other domains (e.g., IT, biology, meteorlogy) can leverage these approaches for builiding their own world models.</p>

<p>LeCun is the closest of them all to my definition, but it remains to be seen if a true, decoupled world model is actually useful for robotics. It’s true that humans have a world model, and many policy failures can be attributed to a lack of a world model. However, it’s probably fair to say that the human world model is part of an “end-to-end model”, and not some modular piece in our brain. Certainly, there is some evidence that world modeling principles are useful training objectives for models (see Meta’s <a href="https://arxiv.org/pdf/2510.02387">Code World Models</a>, <a href="https://arxiv.org/abs/2511.05963">this work</a> combining world modeling with autoregressive language generation, and <a href="https://arxiv.org/pdf/2602.05842">Reinforcement World Model Learning</a> as examples).</p>

<p>If I had to guess, domains that are more easily RL’ed or that can collect enough expert demonstrations can probably learn implicit world models. Other domains will need to learn explicit world models. Explicit world models will help policies generalize to novel scenarios, whereas implicit world models will be easier to train and integrate with policies<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">In gradient descent we trust!</span>
</span>
. Regardless, I’m excited to see where we go from here.</p>]]></content><author><name></name></author><category term="blog" /><summary type="html"><![CDATA[Many AI researchers are building world models, and while we conceptual understand what a world model is, I feel that we lack a technical definition of one. This post gives one.]]></summary></entry><entry><title type="html">A PhD Student’s Honest Reflection on Attending CVPR as a First-Time Author</title><link href="https://skumar-ml.github.io/blog/cvpr-first-time-author-reflection/" rel="alternate" type="text/html" title="A PhD Student’s Honest Reflection on Attending CVPR as a First-Time Author" /><published>2026-06-07T17:00:00+00:00</published><updated>2026-06-07T17:00:00+00:00</updated><id>https://skumar-ml.github.io/blog/cvpr-first-time-author-reflection</id><content type="html" xml:base="https://skumar-ml.github.io/blog/cvpr-first-time-author-reflection/"><![CDATA[<p>CVPR 2026 was my first time attending a conference as an author<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">I also attended CVPR 2024, but not as an author.</span>
</span>
. I went to all five days (two workshop days, three main conference days), and as the days went on, I maintained a Note on my phone with any fleeting thoughts I had. This post compiles &amp; reflects on them.</p>

<figure class="post-figure post-figure--half">
  <img src="/assets/images/cvpr_0.JPG" alt="Shubham Kumar at the CVPR 2026 conference sign" />
  <figcaption>Me at CVPR 2026!</figcaption>
</figure>

<h2 id="the-conference-is-enormous-and-most-of-it-is-noise"><a href="#the-conference-is-enormous-and-most-of-it-is-noise"></a>The conference is enormous, and most of it is noise</h2>

<p>I knew this from going to CVPR in 2024, but the scale of this conference is immense. Naively trying to see everything is a fool’s errand.</p>

<figure class="post-figure post-figure--half">
  <img src="/assets/images/cvpr_busy.jpg" alt="The reception dinner at CVPR" />
  <figcaption>The reception dinner at CVPR (this picture shows maybe 1% of the attendees).</figcaption>
</figure>

<p>It’s important to recognize that, at this scale, most of the conference is noise. By noise, I don’t mean low-quality posters or talks (although, there is a fair share of that too). I mean that there will be about a dozen papers, a handful of people, and maybe one workshop session that are actually high-signal <em>to you</em>. The rest will be irrelevant to your research.</p>

<p>So then, the goal <em>before</em> coming to the conference is to pre-identify them (to the extent possible). Let’s walk through each event type and how to navigate the signal-to-noise ratio.</p>

<h3 id="workshops-talks-and-posters"><a href="#workshops-talks-and-posters"></a>Workshops (Talks and Posters)</h3>
<p>Workshop talks are only worth it if there is a specific speaker you genuinely want to see<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">One notable exception is community-oriented workshops (e.g., <a href="https://sites.google.com/view/bitterlessonscv" target="_blank" rel="noopener noreferrer">Bitter Lessons</a>), where you can learn from interesting experiences from established, well-known researchers.</span>
</span>
. For example, if you’ve closely been following someone’s work, or if your next research idea builds upon the presenter’s past work. Otherwise, to be blunt, most people are bad presenters, and you’ll just be wasting your time. More on this later.</p>

<p><strong>I strongly prefer</strong> going to the workshop poster session. It’s very interactive, and unlike the main conference poster session, it is not crowded. This means that you can get <em>a lot</em> of high-quality time with anyone you want to meet. This is one of the best ways to form connections at a conference.</p>

<p>However, note that many posters (even for a session focused on your research area) may not be interesting <em>to you</em>. You may find that the problem statement is narrow (and sometimes already addressed) or the approach incrementally builds on past methods. You can mostly predict signal depending on the authors (specifically, the PIs) and (sometimes) the affiliations on the work, but it’s important to be open-minded; you never know when a poster will surprise you.</p>

<h3 id="orals"><a href="#orals"></a>Orals</h3>

<p>Orals are generally not worth attending. Not because the ideas are not good; they are. But because there is no control on the quality of the presentation (or presenter), and by design, there’s no interaction. Don’t take my word for it—see what people had to say about this CVPR’s orals:</p>

<div class="post-embed">
<blockquote class="twitter-tweet"><p lang="en" dir="ltr">It’s kind of sad that 90%+ of the oral sessions are just presenters standing on the podium reading script in front of poorly designed slides</p>&mdash; Jia-Bin Huang (@jbhuang0604) <a href="https://x.com/jbhuang0604/status/2063296968710087135?ref_src=twsrc%5Etfw">June 6, 2026</a></blockquote>
<blockquote class="twitter-tweet"><p lang="en" dir="ltr">CVPR oral presentation this year is a disaster. They read scripts, copy-pasted tables from paper and couldn&#39;t answer questions because they arent the people who did the work.<br /><br />CVPR organizer should not let these talk at the precious time of thousands of researchers.</p>&mdash; Viet Lai (@laidacviet) <a href="https://x.com/laidacviet/status/2063300304859447497?ref_src=twsrc%5Etfw">June 6, 2026</a></blockquote>
<script async="" src="https://platform.x.com/widgets.js" charset="utf-8"></script>
</div>

<p>This time might be better spent relaxing, exploring company demos (in case you find that interesting), having fun in the city you are in, or setting up meetings with folks.</p>

<h3 id="main-track-poster-session"><a href="#main-track-poster-session"></a>Main-track Poster Session</h3>

<p>I spent most of my time in the poster session. Like the workshop posters, the main-track posters is where you can have real conversations—but it can get <strong>crowded fast</strong>. If you want to talk to a specific author, go early. You can even offer to help them set up. Going early is a reliable way to get individual (or small group) time with the author before the rush begins (and they will also be more relaxed/fresh).</p>

<p>At CVPR, there were ~650 posters in each poster session. My labmate insisted on walking though every row<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">I told him he was crazy.</span>
</span>
. I opted to pre-select the posters I would be interested in by using <a href="https://www.scholar-inbox.com/">Scholar Inbox</a>. It’s a fantastic software—you should use it. I wish they had an app.</p>

<figure class="post-figure post-figure--half">
  <img src="/assets/images/cvpr_labmate.jpg" alt="Me and my labmate at CVPR badge registration." />
  <figcaption>Me and <a href="https://samyakr99.github.io/" target="_blank" rel="noopener noreferrer">said labmate</a> at CVPR badge registration.</figcaption>
</figure>

<h2 id="presenting-is-an-underrated-skill-that-almost-nobody-has"><a href="#presenting-is-an-underrated-skill-that-almost-nobody-has"></a>Presenting is an underrated skill that almost nobody has</h2>

<p>I mentioned this above, but I wanted to reflect more on it here. One thing that really struck me during this conference: <strong>most people are not good at presenting</strong>. This is true across the spectrum (from PhD students to senior PIs). Getting good at presenting is an underrated skill, and it is also very difficult.</p>

<p>Part of the difficulty is structural. You do not present often enough to build real fluency, and when you do, you rarely get feedback. It’s mostly considered rude for audience members to critique your presentation. Even if you do get signal, one rarely iterates on the presentation and gives it again (i.e., there’s no RL feedback loop). This is all compounded by the fact that presenting is not taught as a fundamental skill to PhD students, and the <em>perceived</em> reward for being good at presenting is not very high<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">You get your PhD degree after you’ve published enough work. You do need to present your work in a defense, but from what I understand, folks are usually not prevented from graduating solely based on their presentation skills.</span>
</span>
.</p>

<h2 id="the-strategy-behind-meeting-people"><a href="#the-strategy-behind-meeting-people"></a>The strategy behind meeting people</h2>

<p>CVPR had (on the order of) 10,000 attendees. Good luck if you want to “cold” meet someone. You have to remember that a lot of these interactions are mediated by status. Assuming you’re like me (a PhD student who is a first-time author), your status is largely determined by your institution and your advisor. The larger the gap in status, the less likely you are to meaningfully network with a person of interest. Nevertheless, there are some specific strategies to increase your chances:</p>

<h3 id="1-reaching-out-beforehand"><a href="#1-reaching-out-beforehand"></a>1. Reaching out beforehand</h3>
<p>A brief, specific email can go a long way. PhD students and recent grads tend to be the most receptive, mainly because they also want to network (and the status disparity is minimal). If the person of interest is a well-known PI, you may not get a response. In such cases, it is often more productive to reach out to their students whose work you find interesting and relevant. But regardless, there’s no harm in shooting your shot<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">I emailed a Goodfire researcher (we’ve had sparse correspondence in the past, and I’ve closely followed their research) with little hope of hearing back. To my surprise, they responded, and we ended up chatting for 45 mins. That connection would not have formed if not for the email.</span>
</span>
.</p>

<h3 id="2-workshop-posters"><a href="#2-workshop-posters"></a>2. Workshop posters</h3>

<p>As mentioned above, workshop posters are a great way to meet people. If you have a good conversation and want to connect with author, exchange contact info<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">This is critical—you need to close the loop in the moment.</span>
</span>
 and message them later to tell them when your poster is. Also, add them on LinkedIn or X.</p>

<h3 id="3-coming-from-a-known-lab-helps"><a href="#3-coming-from-a-known-lab-helps"></a>3. Coming from a known lab helps</h3>

<p>If your lab is known, or if there are many students in your lab, chances are that some of your labmates will be there. They may have connections or friends of their own, who may become your connections and friends. Don’t be shy to introduce yourself to them. Also, your advisor may have alumni who attend. They should  generally be open to meeting. Even if these connections are not super relevant to your area, you should still nurture them. It’s good to make friends, and they will make future conferences feel less foreign.</p>

<figure class="post-figure post-figure--half">
  <img src="/assets/images/cvpr_labAlum.JPG" alt="Some alum from our lab." />
  <figcaption>Meeting up with some alumni from our lab after the reception dinner.</figcaption>
</figure>

<p>For more structured advice on networking, I recommend <a href="https://twitter.com/jbhuang0604/status/1517352789780934656">Jia-Bin Huang’s guide to networking at an in-person conference</a> (and his broader <a href="https://github.com/jbhuang0604/awesome-tips">awesome-tips</a> collection).</p>

<h2 id="pace-yourself--seriously"><a href="#pace-yourself--seriously"></a>Pace yourself — seriously</h2>

<p>You don’t need to see everything. It’s okay to skip oral sessions or to sleep in. Take care of yourself, so that you can get the most value out of the time you do spend at the conference.</p>

<h2 id="the-food-is-bad"><a href="#the-food-is-bad"></a>The food is bad</h2>

<p>I will not belabor this, but it needs to be said. Plan accordingly<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">But my labmate said WACV—which is roughly 5× smaller than CVPR—had steak for lunch, so maybe it’s just a scale thing.</span>
</span>
.</p>

<h2 id="getting-the-most-out-of-poster-presenters"><a href="#getting-the-most-out-of-poster-presenters"></a>Getting the most out of poster presenters</h2>

<p>The best way to engage with the author is to ask questions. The most basic kind are clarification questions, or specific questions about their method. These are good to help you understand their work but typically don’t lead to deeper discussion.</p>

<p>Think about this way: the author has some deeper intuition of the specific problem/method/area, since they are working knee-deep in it. Your job is to probe for that info, which usually doesn’t happen if you keep the discussion hyper-focused on their poster. Some good questions to ask are:</p>

<ol>
  <li>Have you tried/tested X? Did it work? It’s especially good if X was a failed result, since you can ask why they think that was. This helps unearth some intuition.</li>
  <li>How do you conceptually think about a different approach (Y)? You should be able to have a discussion of the strengths &amp; weaknesses compared to their work.</li>
  <li>How do you want to extend this in the future?</li>
  <li>Share an idea you have if you feel like it’s connected. Talk about how their work would be applicable to their idea. Ask for their opinion.</li>
  <li>Offer them a suggestion to improve their work.</li>
</ol>

<h2 id="presenting-my-poster"><a href="#presenting-my-poster"></a>Presenting my poster</h2>

<p>My poster presentation was on the last slot, on the last day. Needless to say, I was wiped before my presentation started, and beyond tired after it ended. A few thoughts on having a good poster session:</p>

<ol>
  <li>Have water for yourself. You will be talking close to non-stop.</li>
  <li>If someone stares at your poster for 5-10 seconds, ask them if they want you to walk through the poster. Don’t just stand there.</li>
  <li>If there are many people around, try to talk loudly (or as loud as you can).</li>
  <li>I always like asking people what they work on. You never know how the conversation will evolve.</li>
  <li>If I spend enough time conversing with someone, I ask them to connect on social media.</li>
</ol>

<figure class="post-figure post-figure--half">
  <img src="/assets/images/cvpr_poster.JPG" alt="My poster!" />
  <figcaption>Just before my poster session starting.</figcaption>
</figure>

<h2 id="to-conclude"><a href="#to-conclude"></a>To conclude</h2>

<p>I strongly believe that you can largely control your own experience. So seize the opportunity of CVPR with both hands, and make the most of it. Talk to people, make connections, and have fun! I feel like conferences are the best part of the PhD experience<span class="footnote" tabindex="0" role="note">
  <sup class="footnote-marker" aria-hidden="true"></sup>
  <span class="footnote-content">Except for graduating, as my labmate reminded me.</span>
</span>
; it surely makes all those hours in the lab worth something.</p>]]></content><author><name></name></author><category term="blog" /><summary type="html"><![CDATA[This post honestly reflects on my CVPR experience, touching on navigating the conference, networking, poster sessions, food, and more!]]></summary></entry><entry><title type="html">Welcome to My Academic Website</title><link href="https://skumar-ml.github.io/blog/welcome-to-my-website/" rel="alternate" type="text/html" title="Welcome to My Academic Website" /><published>2023-01-01T17:00:00+00:00</published><updated>2023-01-01T17:00:00+00:00</updated><id>https://skumar-ml.github.io/blog/welcome-to-my-website</id><content type="html" xml:base="https://skumar-ml.github.io/blog/welcome-to-my-website/"><![CDATA[<p>I’m excited to launch my new academic website! This platform will serve as a hub for my research activities, publications, and teaching materials. I’ll be regularly updating this blog with news about my research, conference presentations, and other academic activities.</p>

<h2 id="what-to-expect"><a href="#what-to-expect"></a>What to Expect</h2>

<p>On this website, you’ll find:</p>

<ul>
  <li>Information about my current research projects</li>
  <li>A complete list of my publications with links to papers</li>
  <li>Details about courses I’m teaching</li>
  <li>Blog posts about my research and academic journey</li>
</ul>

<h2 id="recent-research"><a href="#recent-research"></a>Recent Research</h2>

<p>I’ve been working on [brief description of your current research]. This work aims to [describe the goals and potential impact of your research]. I’m looking forward to sharing more updates as the project progresses.</p>

<h2 id="upcoming-events"><a href="#upcoming-events"></a>Upcoming Events</h2>

<p>I’ll be presenting my research at the following upcoming conferences:</p>

<ul>
  <li><strong>Conference Name</strong> - Date, Location</li>
  <li><strong>Workshop Name</strong> - Date, Location</li>
</ul>

<p>Feel free to reach out if you’re interested in my research or would like to collaborate!</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[I’m excited to launch my new academic website! This platform will serve as a hub for my research activities, publications, and teaching materials. I’ll be regularly updating this blog with news about my research, conference presentations, and other academic activities.]]></summary></entry></feed>