[{"content":" I was recently asked a question which I forgot the intuition for, so I wanted to write a post that approaches it from a few different perspectives and share how I got my intuition back. In the interest of entertaining myself I have reworded the problem with a ridiculous scenario that still preserves the mathematical statement of the problem.\nImagine you get into a high speed chase with the cops. You catch a lucky break and veer out of sight but end up crashing \\(X_\\mathrm{init} = 100\\) meters from the last location where the cops saw you. You proceed to get out of the vehicle and continue running away on foot, but in a state of shock, each stride you either take one step away from the cops with probability \\(p\\) or accidentally take a step backwith probability \\(q = 1 - p\\).\nYour safehouse is just 1,000 meters away from where the cops last saw you (\\(X = 1000\\)), but if you take too many steps back you\u0026rsquo;ll end up being captured by the cops, who are looking for you at your last known location \\((X = 0)\\).\nAssume the distance you cover in a single step is \\(\\Delta x\\) meters and you\u0026rsquo;re physically capable of taking steps as small as \\(0.1\\) meters or as large as \\(100\\) meters (you\u0026rsquo;re either superhuman or you have a collection of ender pearls in your inventory).\nWhat step size should you take assuming you have infinite time, and will only get caught if you stumble to \\(X = 0\\)?\nParameters: $$ \\begin{align*} X_\\mathrm{init} \u0026= 100 \\text{ meters} \\\\ X_\\mathrm{target} \u0026= 1000 \\text{ meters} \\\\ X_\\mathrm{capture} \u0026= 0 \\text{ meters} \\\\ P(\\text{step forward}) \u0026= p \\\\ P(\\text{step backward}) \u0026= q = 1 - p \\\\ \\text{Step size} \u0026= \\Delta x \\in [0.1, 100] \\text{ meters} \\end{align*} $$I\u0026rsquo;ll lead with the intuition and a relatively simple argument first.\nFirst Solution (Quick)\r#\rThe larger steps you take, the more variance you\u0026rsquo;ll experience per step. For example, if you bounce \\(100\\) meters then you have a probability \\(q\\) of getting arrested after just one step. On the other hand, if you take a small step like \\(0.1\\) meters, you only move a little bit, but it\u0026rsquo;ll take you 1,000 steps to cover the same total distance (and I do mean distance in the physics sense of the word, not displacement).\nFor a given \\(\\Delta x\\) we know that the variance of your position increases with the number of steps, but for a given number of steps it decreases when you take smaller steps; so what we\u0026rsquo;re really trying to get at is the tradeoff between having more small steps versus fewer large steps. From the first example we considered it is already apparent that smaller steps are better for the first step, but we can build our intuition more by thinking about what happens when you take more and more smaller steps. Imagine taking the limit \\(\\Delta x \\rightarrow 0\\) for a fixed distance traveled: $$ n \\Delta x = d = \\mathrm{constant} $$As you take smaller and smaller steps you have to take proportionately more steps to move the same distance. Your net displacement is the sum of these steps: $$ \\sum_{i=1}^n \\Delta X_i = n \\overline{\\Delta X} $$In the limit of infinitely many steps \\(\\overline{\\Delta X}\\) converges to the expectation value of a single step: $$ \\overline{\\Delta X} \\rightarrow \\Delta x (p - q) $$Hence your net displacement converges to: $$ n \\Delta x (p - q) = d(p - q) $$which is just a number, not a random variable. In other words, in the limit of infinitely small steps you literally just walk in a straight line, never backtracking towards the police, which is ideal. The closest we can get to that based on the constraints is taking steps of \\(0.1\\) meters, so we choose that as our preferred answer.\nLet\u0026rsquo;s explore what this means from a highly related, but slightly different perspective. To reiterate, a walk is described by your initial position plus the cumulative sum of a sequence of random steps: $$ \\{\\Delta X_i : i = 1, 2, ...\\} $$Each step is IID and can be expressed as: $$ \\Delta X_i = \\Delta x \\, Z_i $$where \\(Z_i\\) is sampled from a distribution that yields \\(+1\\) with probability \\(p\\) or \\(-1\\) with probability \\(q\\). In other words each step has a variance proportional to \\((\\Delta x)^2\\), and the sum of \\(n\\) steps has a variance proportional to: $$ n (\\Delta x)^2 $$Notice that \\(n\\) shows up with a power of \\(1\\) and \\(\\Delta x\\) shows up with a power of \\(2\\). Hence, when we take the limit of small steps, but keep the amount of distance traveled constant we fix a factor of \\(n \\Delta x\\), but the remaining \\(\\Delta x\\) keeps shrinking, collapsing the variance to 0.\nSimulations\r#\rOf course, when I needed it the most, I didn\u0026rsquo;t have the intuition above because I panicked. I was able to come up with the explanation above by doing my favourite past-time activity: writing up Monte Carlo simulations 😀\nI generated 3 simulations of paths taken for each value of \\(dx\\) and plotted them in the figure below, along with the expected value \\(X_n = x_\\mathrm{init} + n (p - q) dx\\). By adjusting the slider at the bottom of the figure you can switch between \\(dx = 0.1, 10\\), and \\(100\\) to see how the step size affects the \u0026lsquo;stability\u0026rsquo; of escape attempts.\nInteractive Monte Carlo Simulations: Effect of Step Size on Escape Probability\rIn-depth solution\r#\rIn this section I\u0026rsquo;ll share a more complicated solution which is how I originally understood this problem, but it\u0026rsquo;s too long and impractical to use in an interview setting unless you happen to memorize the final answer (which I never do).\nLet\u0026rsquo;s forget about the original problem statement for now, and just treat it like a Markov process with absorbing boundary conditions at \\(x_\\mathrm{low} = 0\\) and \\(x_\\mathrm{high} = 1000\\). Fix the step size \\(\\Delta x\\) and define the grid of valid states as \\({x_k = k \\Delta x : k = 1, 2, \\dots N }\\), where \\(N \\equiv x_\\mathrm{high} / \\Delta x\\). As in the original prompt we\u0026rsquo;ll denote the probability of moving towards \\(x_\\mathrm{high}\\) in a given step by \\(p\\), and we\u0026rsquo;ll let \\(q = 1 - p\\). For convenience, denote the probability of escaping the police, starting at state \\(x_m\\) by \\(q_m\\). That is,\n$$ q_m \\equiv \\mathrm{Pr}(\\, \\mathrm{escape} \\, |\\, x_\\mathrm{initial} = x_m \\,) \\ . $$Trivially, at \\(m = 0\\) and \\(m = N\\) we have,\n$$ \\begin{align*} q_0 \u0026= 0 \\\\ q_N \u0026= 1 \\end{align*} $$Doing a \u0026rsquo;next-step\u0026rsquo; analysis, we observe that if you\u0026rsquo;re at state \\(x_m\\) now, then after one step you will either be at \\(x_{m-1}\\) or \\(x_{m+1}\\), at which point if we knew \\(q_{m-1}\\) and \\(q_{m+1}\\) then we could calculate \\(q_m\\) as,\n$$ \\begin{align*} q_m = p \\, q_{m + 1} + q \\, q_{m-1} \\ . \\end{align*} $$Using the fact that \\(p + q = 1\\) we can write the left hand side as\n$$ \\begin{align*} q_m = (p + q) \\, = p \\, q_m + q \\, q_m \\ . \\end{align*} $$Subtracting the latter from the RHS yields\n$$ \\begin{align} 0 = p \\, (q_{m + 1} - q_m) + q \\, (q_{m-1} - q_m) \\end{align} $$Now, if we define \\(F_m \\equiv q_{m} - q_{m - 1}\\) and introduce the inverse odds \\(r \\equiv q / p\\) we can rearrange eqn (1) as\n$$ \\begin{align*} F_{m+1} = \\frac{q}{p} \\, F_m = r F_m \\ . \\end{align*} $$Induction then tells us\n$$ \\begin{align} F_{m+1} = r^m \\, F_1 = r^m q_1 \\ . \\end{align} $$On the other hand, using the definition of \\( F_m \\) we observe that \\( q_m = \\sum_{k=1}^m F_k \\). So, in combination with eqn (2) we have\n$$ \\begin{align*} q_m = \\sum_{k=1}^m F_k = \\sum_{k=1}^m r^{k-1} q_1 = \\frac{q_1}{r} \\sum_{k=0}^{m-1} r^{k-1} = \\frac{q_1}{r}\\frac{1 - r^m}{1 - r}\\ . \\end{align*} $$To solve for \\(q_1\\) we apply the boundary condition \\(q_N = 1\\):\n$$ \\begin{align*} 1 = q_N = \\frac{q_1}{r}\\frac{1 - r^N}{1 - r} \\Rightarrow \\frac{q_1}{r} = \\frac{1 - r}{1 - r^N} \\ . \\end{align*} $$And so,\n$$ \\begin{align} \\boxed{q_m = \\frac{1 - (q / p)^m}{1 - (q / p)^N}}\\ . \\end{align} $$Eqn (3) can be used to solve our original problem. Since \\(N = x_\\mathrm{end} / \\Delta x\\) and \\(m = x_\\mathrm{initial} / \\Delta x\\) we immediately see that taking smaller steps is equivalent to sending $m$ and $N$ to infinity at the same rate, so\n$$ \\begin{align*} \\lim_{\\Delta x \\rightarrow 0} \\mathrm{Pr}(\\mathrm{escape}) = 1. \\end{align*} $$And clearly we should pick the smallest possible value of \\( \\Delta x \\) permitted.\n","date":"3 October 2025","externalUrl":null,"permalink":"/posts/gamblers-ruin/","section":"Posts","summary":"","title":"Drunken Walk / Gambler's Ruin","type":"posts"},{"content":" Ray Hagimoto Physics PhD\nRice University\n\u003c?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"no\"?\u003e Education PhD in Physics (Dec 2024)\nRice University\nBSc in Physics (May 2016)\nUniversity of Texas at San Antonio\nBiography\r#\rAcademic background\r#\rI completed my PhD in Physics at Rice University in Houston. I was advised by Professors Andrew Long and Mustafa Amin. My research consists of quantifying cosmological signatures of axion-like particles, looking for evidence of them in cosmological and astrophysical data. In particular, I focused on the prospect of detecting axion strings via the cosmic microwave background using an effect known as cosmic birefringence. A full list of my publications can be found here, and a copy of my thesis is available here.\nAs an undergrad I did research in computational nanophotonics with Nicolas Large at the University of Texas at San Antonio and did NSF REUs at Brigham Young University and the University of Chicago where I studied quantum dynamics and cosmology respectively.\nA copy of my resume can be found here.\nPersonal background\r#\rMy parents are Japanese and Spanish/American. I grew up in Singapore and moved to the Texas for university, where I\u0026rsquo;ve lived ever since.\nOutside of physics I enjoy origami, photography, and cooking!\nRecent\rDrunken Walk / Gambler's Ruin\r3 October 2025\u0026middot;1456 words\rBuilding a real-time wildlife detection and alert system on a budget\r21 June 2025\u0026middot;2559 words\rCreating a WSL Context Menu in Windows 11\r20 June 2025\u0026middot;994 words\rLinear regression in a nutshell\r3 June 2025\u0026middot;1332 words\rA cute linear regression brainteaser\r31 December 2024\u0026middot;1117 words\rA quant probability question and physics\r8 December 2024\u0026middot;501 words\rShow More\r","date":"3 October 2025","externalUrl":null,"permalink":"/","section":"Home","summary":"","title":"Home","type":"page"},{"content":"","date":"3 October 2025","externalUrl":null,"permalink":"/posts/","section":"Posts","summary":"","title":"Posts","type":"posts"},{"content":"This blog post has an accompanying iPython notebook.\nRecently I went to visit my Mum and I was amazed at the amount of wildlife that visits the backyard. Here\u0026rsquo;s a sample of a few critters that I\u0026rsquo;ve managed to photograph which includes a groundhog, a deer, some birds, and a bee: Previous\rNext\rOne morning my mum mentioned to me that there\u0026rsquo;s a particular bird that she sees occasionally which she\u0026rsquo;d like to get a close-up of to identify it, so I figured that since I have a camera and a tripod I would give it a shot. I decided my best bet was to point my camera at one of the bird baths we have in the backyard and set it to take pictures periodically for a few hours and check the results.\nI had considered building a motion-activated remote control system but I found that using an external device to control my camera was too unstable over WiFi or USB so I figured it was just easier to have the camera take pictures every 2 seconds and review the camera roll after.\nOne photo every 2 seconds is 1,800 photos an hour \u0026ndash; most of which (about 97.5%) are beautifully sharp images of an empty bird bath, e.g.\nRapidly flicking through the images is only marginally more interesting than watching paint dry because you at least get to see the ambient light changing and the leaves rustling, but even so, after flicking through the first 300 photos hoping to spot a bird, I found nothing interesting. I realized that I didn\u0026rsquo;t want to spend the rest of my life holding my pinky down on the right arrow key and praying I didn\u0026rsquo;t miss a bird, so I decided to automate the process.\nInitial thoughts\r#\rInitially my goal was to simply build a computer vision model to locally identify which photos contained birds and which ones didn\u0026rsquo;t, but I quickly became more ambitious. For one thing, when I leave my camera outside to take pictures I have to wait for the timelapse session to end before I can review the pictures and find out if I captured any pictures of birds, which simply isn\u0026rsquo;t good enough. I wanted LIVE updates of bird-images to see the birds in real time. As an example, here\u0026rsquo;s what the notifications look like on my phone:\nRather than describing the different things I tried I\u0026rsquo;ll start by discussing what I ended up building, pointing out the rationale behind some of the design choices. I ended up building two computer vision systems.\nProject Constraints and Budget\r#\rBefore diving into the technical details, I should mention the constraints I set for this project:\nHardware Constraints:\nI can only use existing hardware that I have (my camera, my tripod, my old Android phone) Budget Constraints: 2. Spend as little money as possible\nOutcomes:\nI spent $3.49 on the Tasker app (for Android automation) AWS Lambda has a generous free tier! AWS S3 also has a generous free tier, so I didn\u0026rsquo;t spend any money here These constraints actually forced me to be more creative with my solution. Instead of buying specialized wildlife cameras or motion sensors, I had to work with what I had and leverage cloud services that offer free tiers for small-scale projects like this.\nConvolutional Autoencoder (offline, no alert system, more reliable)\r#\rThe most reliable system I came up with to quickly and accurately detect anomalous (bird) images is a convolutional autoencoder. The basic idea is to imagine we have a black box which could take any image and convert it to a \u0026ldquo;pure background\u0026rdquo;. For example, if the input image is a pure background, it is just the identity transformation. However, if the input image contains a bird, it would return the same image with the bird removed (like using generative fill in Photoshop). I figured that since birds show up so infrequently (less than 3% of the images), I could train an autoencoder on the entire image dataset (including the anomalous images). Here\u0026rsquo;s an example of how it works on a picture that contains a bird:\nI found that a convolutional autoencoder with 57,037 tunable parameters was sufficiently good at doing background reconstruction that I was able to achieve a false positive rate of ~10% and a false negative rate of 0%.\nAutoencoder anomaly detection pipeline in more detail\r#\rTo explain the pipeline in a bit more detail. Here\u0026rsquo;s what I did:\nTrain the autoencoder on background patterns: I split my dataset into 60% for training and 10% for validation, using approximately 2,800 images total (out of 4,000). The training takes about 4 minutes on my GPU. I used a log-MSE loss function, which applies a logarithmic transformation to the mean squared error between the original and reconstructed images. I found that whether I used a log-MSE or MSE I got similar performance, so the log-MSE choice is more because I tried it out and was too lazy to change it back after I found it performed the same.\nCompute localized anomaly scores: For each image, I calculate a specialized \u0026ldquo;localized reconstruction score\u0026rdquo;, which is different from the training loss. While the training loss measures global reconstruction quality across the entire image, this score is designed to identify local clusters of large differences. Here\u0026rsquo;s how it works: First, I compute the pixel-wise difference between the original image and the autoencoder\u0026rsquo;s reconstruction. Then I apply Gaussian smoothing to reduce noise and make local peaks more prominent. Next, I progressively downsample this smoothed difference map using max pooling until it reaches 4×4 resolution, keeping only the maximum value in each region. Finally, I take the highest value from this 4×4 map as the anomaly score. This approach is much better at detecting localized anomalies (like birds) than global metrics because it focuses on the most significant local reconstruction errors rather than averaging across the entire image. This step takes about 10 seconds.\nDetermine the anomaly threshold using Otsu\u0026rsquo;s method: After computing scores for all images in the training and validation sets, I use Otsu\u0026rsquo;s method to automatically determine the optimal threshold for separating normal from anomalous images. This method finds the threshold that maximizes the separation between two classes by analyzing the histogram of scores. It\u0026rsquo;s much more reliable than manually setting a threshold and adapts automatically to the specific characteristics of your dataset. This step is almost instant (~5 milliseconds).\nApply the threshold to find anomalies: With the threshold determined, I can now classify any image as anomalous if its score exceeds this threshold. Since I\u0026rsquo;ve already computed scores for all images, this classification step is nearly instantaneous ~5 milliseconds. The final output is a list of filenames corresponding to images that likely contain birds.\nWhat did the autoencoder learn?\r#\rThis was probably my favourite part of the project. I wanted to visualize how the decoder was mapping from the 2D latent space to images so I created an interactive figure with two axes (shown below). On the right is a depiction of the latent space, with a 2D heatmap highlighting points in the latent space where my images were mapped to. The histogram is bounded by a red box, outside which none of the images in my dataset were mapped to. (So as far as the decoder is concerned, it\u0026rsquo;s the Wild West). On the left is the image computed by feeding the latent space point through my decoder. By clicking different points in the latent space I was very satisfied to find that the autoencoder had indeed learned many different lighting configurations of the bird bath background. In the animation below you can see me clicking and dragging the red dot to visualize different parts of the latent space. As the red dot moves around you can see the decoder smoothly interpolating between different angles of sunlight, with different intensities. The smooth interpolation is particularly compelling. Moreover, I noticed that even when I took the red dot out of the bounding box (especially in the far bottom-left region) the decoded image still looks like a reasonable extrapolation.\nInteractive visualization of the autoencoder's latent space Why are false positives happening?\r#\rThere are some instances when a picture will have a localized reconstruction score that exceeds the threshold set by Otsu\u0026rsquo;s method. Before I explain Otsu\u0026rsquo;s method in a bit more detail it\u0026rsquo;s instructive to look at the distribution of scores to see why Otsu\u0026rsquo;s method might be suited to this problem. Below is a histogram of the reconstruction scores on a log-log scale. Notice that this distribution is clearly bimodal, with a clear separation between the background class (left) and the anomaly class (right) \u0026ndash; this is one of the benefits of choosing a localized difference score; I found that when I experimented with a global difference score the distributions had far more overlap.\nOtu\u0026rsquo;s method systematically scans different thresholds for the binary classification e.g.\n$$ \\begin{align*} \\mathrm{score} \u003e \\mathrm{threshold} \\Rightarrow \\mathrm{anomaly} \\end{align*} $$and chooses the threshold which minimizes the weighted averaged of intra-class variances for the background and anomaly classes respectively (where the weights are given by the class probabilities).\nThe plot above shows that the threshold computed using Otsu\u0026rsquo;s method (black dashed line) is in the far right tail of the distribution of background scores, but not quite in the empty space between the two distributions. I think this is actually ideal because I want my anomaly detector to have as few false negatives as possible, and that means having a higher tolerance for false positives. I\u0026rsquo;m okay with seeing a few background images because it doesn\u0026rsquo;t take long to manually filter them out. On the other hand, false negatives could mean missing out on the photo of a life time!\nStill, I was curious about why some background images have a reconstruction score about 3 times higher than the modal score.\nFalse positives typically occur at transitions between two drastically different lighting conditions. This figure demonstrates a classic false positive case. On the left is a still image of what appears to be just the background - no bird in sight - yet it was flagged as an anomaly by the autoencoder. When I investigated this particular image, I looked at the photos taken immediately before and after it and discovered that this image was captured during a sudden transition between two distinct lighting conditions. Perhaps a cloud had parted briefly, allowing direct sunlight to hit the bird bath for just a moment, or maybe the sun emerged from behind a tree. The autoencoder, having been trained primarily on more stable lighting conditions, struggled to reconstruct this transitional state and flagged it as anomalous.\nThe right side of the figure shows an animation that sequences through the images before, during, and after this lighting transition. The animation pauses when it reaches the problematic image (indicated by the play button overlay) so you can see exactly when in the sequence the false positive occurred. This visualization helps explain why the autoencoder was confused - the sudden change in lighting created a pattern it hadn\u0026rsquo;t seen enough of during training.\nReal-time system\r#\rThe goal of the real-time system is to enable my phone to receive alerts when my camera has taken pictures of a bird within seconds of the picture being taken, and without me having to go outside and interrupt the timelapse.\nIn order to get this to work I had to use a different computer vision model which is more lightweight than the PyTorch Autoencoder. The resulting model is less reliable but it\u0026rsquo;s fast enough to work in real-time on AWS Lambda, even accounting for cold starts. In my stress testing I found that it can handle the fastest setting for my camera\u0026rsquo;s timelapse interval which is 1 photo per second.\nBelow is a picture of the physical set up: my camera on a tripod, pointing at the bird bath, with a USB-C cable connecting it to my phone. The phone is then able to handle HTTP requests and communicate with my cloud-deployed computer vision model.\nTo be more precise, the way it works is that I connect my camera to my phone via USB-C and have my camera write two files (i) a raw (24 MP .ARW files) saved to an internal SD card and (ii) a low quality (2-7 MB) .jpg to my Phone\u0026rsquo;s storage.\nHere\u0026rsquo;s how the cloud pipeline works:\nDual file capture: My camera simultaneously saves high-quality raw files (24 MP .ARW) to an internal SD card and lower-quality JPEGs (2-7 MB) to my phone\u0026rsquo;s storage via USB-C connection.\nFile detection: On my phone, I use the Tasker app to monitor for new .jpg file creation events in my Pictures folder.\nSecure upload preparation: When a new image is detected, Tasker triggers an HTTP request to my AWS Lambda function, which generates a pre-signed URL for secure S3 upload without requiring AWS credentials on my phone.\nCloud upload: The pre-signed URL allows temporary upload permissions to my S3 bucket, where the image is stored securely.\nAutomated analysis: S3 bucket events trigger my computer vision Lambda function to analyze the uploaded image using the background subtraction algorithm.\nReal-time notification: If an anomaly (hopefully a bird) is detected, the system sends the image to my Telegram account via a bot I configured.\nThe system is designed to handle images uploaded once per second (the maximum rate my camera can achieve) and can process them efficiently even with AWS Lambda cold starts.\nComputer vision model\r#\rThis technique also uses a form of background subtraction, but instead of using an autoencoder to produce pure background images, it uses an exponential moving average of the photos over the last 30 seconds. I found 30 seconds to be a good time frame because it prevents false positives from being caused by gradual changes in lighting.\nHere\u0026rsquo;s how the real-time detection works:\nBackground modeling: For each new image, I update an exponential moving average (EMA) of the background using a decay factor of 0.05. This creates a smooth, adaptive background model that gradually adapts to lighting changes.\nLuminance comparison: I convert both the current image and background model to grayscale using standard luminance coefficients (0.2126×R + 0.7152×G + 0.0722×B), then compute the absolute difference between them.\nNoise reduction and thresholding: I apply Gaussian blur to the difference map to reduce noise, then threshold it to create a binary map highlighting regions with significant differences from the background.\nContour analysis: I find contours in the binary map and filter them based on several criteria: area (between 0.15% and 5% of image), aspect ratio (max 2:1), vertical position (preferring middle 70% of image), and contrast (minimum 10 gray levels).\nScoring and detection: Each valid contour is scored based on its area and the mean squared luminance difference within the contour. The highest-scoring contour above a minimum threshold triggers an anomaly detection.\nThe system requires at least 10 observations before making predictions to ensure the background model has stabilized. This approach is much faster than neural networks and can run efficiently on AWS Lambda with cold starts.\nSuccess!\r#\rAnyways, after all this I eventually got lucky and snapped a picture of the bird my Mum was curious about. Turns out it is a male House finch!\nHere are some more images to enjoy\nPrevious\rNext\r","date":"21 June 2025","externalUrl":null,"permalink":"/posts/bird-watching-anomaly-detection/","section":"Posts","summary":"","title":"Building a real-time wildlife detection and alert system on a budget","type":"posts"},{"content":" How to add an Open WSL here context menu in Windows 11\r#\rIn this blog post I will explain how I added an \u0026lsquo;Open WSL here\u0026rsquo; context menu in Windows 11.\nScreenshot of Windows Explorer \u0026lsquo;Open WSL here\u0026rsquo; context menu option\rThis solution opens Windows Terminal with the correct terminal profile so that the tab uses the Ubuntu icon and tab name.\nTL;DR Solution:\r#\rFind your wt.exe path\n(Get-Command wt.exe).Source\rFor me it returned something like:\nC:\\Users\\ray\\AppData\\Local\\Microsoft\\WindowsApps\\wt.exe\rCreate the per-user registry key\nOpen Regedit.\nGo to (or create) these keys under your user hive (HKEY_CURRENT_USER):\nHKEY_CURRENT_USER\\Software\\Classes\\Directory\\Background\\shell\\WSL\rUnder WSL, make a command sub-key.\nSelect command, double-click (Default), and set it to:\n\u0026#34;C:\\Users\\ray\\AppData\\Local\\Microsoft\\WindowsApps\\wt.exe\u0026#34; -p \u0026#34;Ubuntu 24.04.1 LTS\u0026#34; -d %V\rClose Regedit, then restart Explorer or log out and back in.\nGet WSL Profile name\nOpen Windows Terminal in Powershell or Command Prompt Run wsl --list Identify the profile you want to use (for me it\u0026rsquo;s Ubuntu-24.04) Write it down or remember it. Update Ubuntu Profile \u0026ldquo;Command Line\u0026rdquo; Setting\nOpen Windows Terminal. Go to the settings panel (CTRL + ,) or Dropdown Menu -\u0026gt; Settings. Click on the dropdown arrow next to \u0026ldquo;Command line\u0026rdquo;. Change it from whatever it currently is to C:\\Windows\\system32\\wsl.exe -d \u0026lt;NAME_FROM_STEP_3\u0026gt;\r(For me it was C:\\Windows\\system32\\wsl.exe -d Ubuntu-24.04). Update \u0026ldquo;Starting Directory\u0026rdquo; Setting\nRight underneath the \u0026ldquo;Command line\u0026rdquo; option we just set, you will see \u0026ldquo;Starting directory\u0026rdquo;. If you don\u0026rsquo;t do this then when you open WSL in a new tab from Windows terminal it\u0026rsquo;ll always start in system32/ . Change the \u0026ldquo;Starting directory\u0026rdquo; option to \u0026ldquo;~\u0026rdquo;.\nI was worried that this would break some WSL behaviors like inheriting the directory from the process that launches it but so far I haven\u0026rsquo;t run into any undesired behavior. Things I\u0026rsquo;ve tested: Opening WSL from Powershell opens WSL at whatever directory Powershell was in; opening WSL using the \u0026ldquo;Open WSL Here\u0026rdquo; context menu option opens WSL at that directory. (Optional) Add an icon\nNavigate to the same key we in we used in step 2.2:\nHKEY_CURRENT_USER\\Software\\Classes\\Directory\\Background\\shell\\WSL\rRight-click the right panel -\u0026gt; New -\u0026gt; String Value\nName it \u0026ldquo;Icon\u0026rdquo;\nRight-click on \u0026ldquo;Icon\u0026rdquo; and select \u0026ldquo;Modify\u0026rdquo;\nEnter the path to your desired .ico. E.g. I made a directory at C:\\Users\\\u0026lt;USERNAME\u0026gt;\\WindowsTerminalIcons containing a file called ubuntu.ico which I generated by finding a .png of the Ubuntu logo and then finding a png to ico converter online. So I entered C:\\Users\\\u0026lt;USERNAME\u0026gt;\\WindowsTerminalIcons\\ubuntu.ico.\nThings I tried / common pitfalls\r#\rEnvironment details\r#\rTo give some context: I\u0026rsquo;m on Windows 11, and I installed Windows Terminal from the Microsoft Store. That means wt.exe is really just an alias for:\nC:\\Users\\\u0026lt;you\u0026gt;\\AppData\\Local\\Microsoft\\WindowsApps\\wt.exe\rIf I had installed Windows Terminal by downloading the executable installer the traditional way from the Microsoft website, I\u0026rsquo;d have a real binary at C:\\Program Files\\Windows Terminal\\wt.exe and none of the alias drama below would have been an issue.\nTip: Before beginning, check if you can open Windows Terminal by pressing Win + R, typing wt.exe, and hitting Enter. If it opens you don\u0026rsquo;t need to provide the full path to the wt.exe executable when you write your registry key command, unlike me.\nWhy I worked under HKEY_CURRENT_USER instead of HKEY_CLASSES_ROOT\r#\rI first tried editing the key I found in the screenshot under:\nHKEY_CLASSES_ROOT\\Directory\\Background\\shell\\WSL\\command\rBut I got an \u0026ldquo;access denied\u0026rdquo; message. By consulting ChatGPT, I learned that anything I put under:\nHKEY_CURRENT_USER\\Software\\Classes\\Directory\\Background\\shell\\WSL\rwould override the global settings in HKEY_CLASSES_ROOT. And since that HKEY_CURRENT_USER path belongs to me, I could edit it without admin rights.\nPicking the right placeholder: %V vs %1\r#\rWhen Explorer runs my command, it replaces:\n%V → the folder path when I right-click the background of a folder window %1 → the path when I right-click the folder icon itself I wanted the menu item on the empty space (background), so I used %V to pass the directory path into my wt.exe call.\nThe console host vs Windows Terminal\r#\rMy first working command was:\nwsl.exe -d Ubuntu-24.04 --cd \u0026#34;%V\u0026#34;\rThat dropped me into Ubuntu-24.04 correctly, but in the old console host, so the tab title and icon as the command prompt\u0026rsquo;s, which looks ugly and that really bothered me. This behaviour is due to the fact that the command is being run from command prompt when you run wsl. Here\u0026rsquo;s an example of what happens when you open WSL from Powershell, showing the same behaviour (notice that the tab name and icon stay as the powershell tab name and icon) Next, I changed it to:\nwt.exe -p \u0026#34;Ubuntu 24.04.1 LTS\u0026#34; -d \u0026#34;%V\u0026#34;\rFrom PowerShell or cmd, that launched Windows Terminal with the right tab title. But when I clicked the context menu entry, nothing happened. No error, no window—just silence. This confirmed that the Store alias wt.exe wasn\u0026rsquo;t resolving in Explorer.\nQuick experiments\r#\rcmd.exe /c start \u0026quot;\u0026quot; wt ... worked but flashed a brief cmd window. Direct wt.exe worked in shells but hung or failed in Explorer and Run. Using the real binary path fixed everything—no flashes, no hangs. Final fix: use the real WT binary\r#\rI ran (Get-Command wt.exe).Source to get the full path to the real EXE, then updated my HKEY_CURRENT_USER registry entry to:\n\u0026#34;C:\\Users\\ray\\AppData\\Local\\Microsoft\\WindowsApps\\wt.exe\u0026#34; -p \u0026#34;Ubuntu 24.04.1 LTS\u0026#34; -d \u0026#34;%V\u0026#34;\rAfter restarting Explorer, Shift + Right-Click on any folder background and choosing Open WSL here fires up Windows Terminal directly into Ubuntu-24.04, with the proper tab title and no extra console windows, but at the wrong directory! I found that by changing the \u0026ldquo;command line\u0026rdquo; setting from ubuntu2404.exe to C:\\Windows\\system32\\wsl.exe -d Ubuntu-24.04 I could solve the issue. I guess the problem is with accessing the ubuntu2404.exe executable directly instead of going through wsl.exe as a middle-man.\nSummary highlights:\nCheck your alias: Win + R → wt.exe Override global registry by editing HKEY_CURRENT_USER\\Software\\Classes Use %V for background clicks Find the real wt.exe path with (Get-Command wt.exe).Source Point your context-menu command at that full path for a clean launch Change \u0026lsquo;command line\u0026rsquo; option from your Linux executable to C:\\Windows\\system32\\wsl.exe -d ProfileName ","date":"20 June 2025","externalUrl":null,"permalink":"/posts/creating-a-wsl-context-menu-in-win11/","section":"Posts","summary":"","title":"Creating a WSL Context Menu in Windows 11","type":"posts"},{"content":" Introduction\r#\rThese notes are written as a quick reference on ordinary least squares (OLS) regression. I have more detailed notes in pdf form here .\nSuppose we have a dataset \\(\\mathcal{D} = \\{(x_i, y_i) : i = 1, 2, \\dots, n\\}\\), where \\( x_i \\in \\mathbb{R} \\) is a scalar covariate and \\( y_i \\in \\mathbb{R} \\) is the response. We assume this data is drawn i.i.d. from some unknown distribution \\( p(x, y) \\). Since we’re interested in predicting \\( y \\) from \\( x \\), it’s natural to factor the joint distribution as:\n\\[ p(x, y) = p(y \\mid x)\\, p(x) \\]In linear regression, we posit a parametric model for the conditional distribution:\n\\[ p(y \\mid x; \\beta) \\propto \\exp\\left( -\\frac{1}{2\\sigma^2} [y - f(x; \\beta)]^2 \\right) \\]where \\( f(x; \\beta) = \\beta_0 + \\beta_1 x \\) and \\( \\sigma^2 \\) is a fixed noise variance. This is equivalent to the model:\n\\[ y = \\beta_0 + \\beta_1 x + \\varepsilon \\;, \\quad \\varepsilon \\sim \\mathcal{N}(0, \\sigma^2) \\]The goal is to find parameters \\( \\beta_0 \\) and \\( \\beta_1 \\) that minimize the squared residuals. Define the loss function:\n\\[ L(\\beta_0, \\beta_1) = \\sum_{i=1}^n \\left(y_i - \\beta_0 - \\beta_1 x_i\\right)^2 \\]Minimizing \\( L \\) with respect to \\( \\beta_0 \\) and \\( \\beta_1 \\) yields the closed-form solution:\n\\[ \\hat{\\beta}_1 = \\frac{\\sum_i (x_i - \\bar{x})(y_i - \\bar{y})}{\\sum_i (x_i - \\bar{x})^2} \\quad \\text{and} \\quad \\hat{\\beta}_0 = \\bar{y} - \\hat{\\beta}_1 \\bar{x} \\]where \\( \\bar{x} = \\frac{1}{n} \\sum_i x_i \\) and \\( \\bar{y} = \\frac{1}{n} \\sum_i y_i \\).\nMultivariate Case\r#\rNow consider the general multivariate case where each input is a vector \\( \\mathbf{x}_i \\in \\mathbb{R}^p \\). To include the intercept \\( \\beta_0 \\), we prepend a 1 to each input vector. That is, we define the augmented input vector \\( \\tilde{\\mathbf{x}}_i = [1, x_{i1}, x_{i2}, \\dots, x_{ip}]^\\mathsf{T} \\in \\mathbb{R}^{p+1} \\).\nStack the data into matrices:\nLet \\( X \\in \\mathbb{R}^{n \\times (p+1)} \\) be the design matrix with rows \\( \\tilde{\\mathbf{x}}_i^\\mathsf{T} \\), Let \\( Y \\in \\mathbb{R}^{n} \\) be the response vector. The model is:\n\\[ Y = X \\beta + \\varepsilon \\;, \\quad \\varepsilon \\sim \\mathcal{N}(0, \\sigma^2 I) \\]where \\( \\beta \\in \\mathbb{R}^{p+1} \\) includes the intercept term as its first entry.\nMaximizing the likelihood (or equivalently minimizing squared error) yields the OLS estimator:\n\\[ \\hat{\\beta}_\\mathrm{OLS} = (X^\\mathsf{T} X)^{-1} X^\\mathsf{T} Y \\]This formula generalizes the scalar case and provides a fast, closed-form solution when \\( X^\\mathsf{T} X \\) is invertible. We will discuss more about when \\(X^\\mathsf{T} X\\) is invertible in the section on multicollinearity.\nAssumptions\r#\rBefore we apply the OLS estimator, it\u0026rsquo;s worth reviewing the key assumptions that underpin it. These assumptions aren\u0026rsquo;t always strictly satisfied in practice — in fact, they’re often violated to some degree. But knowing what assumptions we\u0026rsquo;re making helps us understand when and why the OLS results might be misleading.\nOLS Regression Assumptions\r#\rLinearity Random sampling No perfect multicollinearity No weak exogeneity Homoscedasticity Uncorrelated errors Errors follow a distribution Model specification Linearity\r#\rThis means that the response \\(y\\) is linear wrt the parameters \\(\\beta\\). For example, if we had two covariates \\(x_1\\) and \\(x_2\\), we are assuming that \\(y = \\beta_0 + \\beta_1 x_1 + \\beta_2 x_2\\). The covariates may themselves be nonlinear functions of an underlying variable, e.g. we could model \\(y = \\beta_0 + \\beta_1 x^2\\).\nRandom sampling\r#\rThe data samples \\((x_i, y_i)\\) are assumed to be IID.\nHomoscedasticity\r#\rResiduals have constant variance: \\(\\mathrm{Var}(\\varepsilon_i) = \\sigma^2\\) for all \\(i\\).\nUncorrelated errors\r#\rResiduals are not autocorrelated. That is, they satisfy\n\\(\\mathrm{Cov}(\\varepsilon_i, \\varepsilon_j) = 0\\) for \\(i \\ne j\\).\nNo perfect multicollinearity\r#\rIf you have multiple covariates, perfect multicollinearity means that a subset of the covariates are linearly dependent. The simplest example would be \\(x_1 = x_2\\). More generally, it means there exists a relationship like \\(x_{m_1} + x_{m_2} + \\cdots + x_{m_s} = 0\\), where \\(m_i\\) are the indices of the \\(s\\) multicollinear covariates. Although the assumption is that there is no perfect multicollinearity, even if there is strong multicollinearity OLS regression becomes unreliable, even if the asymptotic properties still hold.\nErrors follow a distribution\r#\rThis is optional, but we typically assume that the residuals are i.i.d Gaussian:\n\\(\\varepsilon \\sim N(0, \\sigma^2)\\).\nModel specification\r#\rThis is a very subtle and easy-to-forget assumption, which basically says that we assume the model is true. In other words, not only do we assume that the true relationship is linear, but also that the only covariates are precisely the ones we included in our formula. This means that we’re assuming there are no other omitted covariates that influence the response (I\u0026rsquo;m using the word \u0026ldquo;cause\u0026rdquo; very loosely here).\nSignificance Testing\r#\rAfter fitting the OLS model, we often want to assess whether each coefficient \\( \\beta_j \\) is statistically different from zero. This is typically done using a t-test, which tests the null hypothesis \\( H_0: \\beta_j = 0 \\) against the alternative \\( H_1: \\beta_j \\ne 0 \\).\nGeneral form\r#\rThe test statistic for coefficient \\( \\hat{\\beta}_j \\) is given by:\n$$ t_j = \\frac{ \\hat{\\beta}_j }{ \\widehat{\\mathrm{SE}}(\\hat{\\beta}_j) } $$where \\( \\widehat{\\mathrm{SE}}(\\hat{\\beta}_j) \\) is the estimated standard error of \\( \\hat{\\beta}_j \\). Under the null hypothesis, and assuming the OLS assumptions hold (particularly homoscedasticity and Gaussian errors), this statistic approximately follows a Student’s t-distribution with \\( n - p - 1 \\) degrees of freedom, where \\( p \\) is the number of covariates. (For example if our model is \\( y = \\beta_0 + \\beta_1 x \\) then \\(p = 1\\)).\nUnivariate case\r#\rFor the simple regression model \\( y = \\beta_0 + \\beta_1 x + \\varepsilon \\), we define the residual variance:\n$$ \\hat{\\sigma}^2 = \\frac{1}{n - 2} \\sum_{i=1}^n \\left( y_i - \\hat{y}_i \\right)^2 = \\frac{1}{n - 2} \\sum_{i=1}^n \\left( y_i - \\hat{\\beta}_0 - \\hat{\\beta}_1 x_i \\right)^2 $$Then the standard error of the slope estimator \\( \\hat{\\beta}_1 \\) is:\n$$ \\widehat{\\mathrm{SE}}(\\hat{\\beta}_1) = \\sqrt{ \\frac{ \\hat{\\sigma}^2 }{ \\sum_i (x_i - \\bar{x})^2 } } $$and the t-statistic becomes:\n$$ t_1 = \\frac{ \\hat{\\beta}_1 }{ \\widehat{\\mathrm{SE}}(\\hat{\\beta}_1) } $$The corresponding p-value is computed from the t-distribution with \\( n - 2 \\) degrees of freedom.\nThe standard error of the intercept \\( \\hat{\\beta}_0 \\) is:\n$$ \\widehat{\\mathrm{SE}}(\\hat{\\beta}_0) = \\sqrt{ \\hat{\\sigma}^2 \\left( \\frac{1}{n} + \\frac{ \\bar{x}^2 }{ \\sum_i (x_i - \\bar{x})^2 } \\right) } \\; . $$ Multivariate case\r#\rIn the multivariate case, the estimated covariance matrix of \\( \\hat{\\beta}_\\mathrm{OLS} \\) is:\n$$ \\widehat{\\mathrm{Cov}}(\\hat{\\beta}) = \\hat{\\sigma}^2 (X^\\mathsf{T} X)^{-1} $$where \\( \\hat{\\sigma}^2 = \\frac{1}{n - p - 1} \\| Y - X \\hat{\\beta} \\|^2 \\) is the residual variance.\nThen for the \\( j \\)-th coefficient, the standard error is:\n$$ \\widehat{\\mathrm{SE}}(\\hat{\\beta}_j) = \\sqrt{ \\left[ \\widehat{\\mathrm{Cov}}(\\hat{\\beta}) \\right]_{jj} } $$and the t-statistic is:\n$$ t_j = \\frac{ \\hat{\\beta}_j }{ \\widehat{\\mathrm{SE}}(\\hat{\\beta}_j) } $$The derivation for these formulas is really easy in the multivariate case.\nPractical use and limitations\r#\rIn practice, these t-tests are used to assess the individual relevance of each covariate. A small p-value (e.g., \u0026lt; 0.05) suggests evidence against the null \\( \\beta_j = 0 \\), implying the corresponding feature may have predictive value.\nHowever, there are important caveats:\nMultiple comparisons: If many covariates are tested, some will appear significant by chance. Adjustments (e.g. Bonferroni correction) may be needed. Model misspecification: If the model omits relevant variables or includes irrelevant ones, the significance test can be misleading. Multicollinearity: High correlation between features inflates standard errors, making it harder to detect significance even when the variable matters. Non-Gaussian errors: The test assumes normally distributed residuals. When this fails, especially in small samples, the p-values may be unreliable. For these reasons, t-tests are often viewed as a rough diagnostic rather than a definitive decision tool. In applied settings (including finance), it is common to combine statistical significance with out-of-sample validation and domain knowledge.\nViolating the assumptions\r#\rWhat happens when the assumptions are violated?\nHeteroscedasticity\r#\rWhen the error variance \\(\\sigma^2\\) is not constant the OLS estimator becomes inefficient\n","date":"3 June 2025","externalUrl":null,"permalink":"/posts/linear-regression-in-a-nutshell/","section":"Posts","summary":"","title":"Linear regression in a nutshell","type":"posts"},{"content":"","date":"30 May 2025","externalUrl":null,"permalink":"/gallery/","section":"Home","summary":"","title":"photography","type":"page"},{"content":" I heard about this problem more than a year ago, I couldn\u0026rsquo;t solve it then and I stopped thinking about it, but I refused to look up a solution because I thought I could do it. Now that I\u0026rsquo;ve finished my PhD I decided to revisit it and finally have an answer :)\nProblem Statement\r#\rConsider the (continuously infinite) set of points \\((x,y)\\) in the interior of a rectangle with length \\(2L\\) and width \\(2W\\) centered at the origin, and whose long side makes an angle of \\(0 \\leq \\theta \\leq \\pi/4\\) with the \\(x\\)-axis. This setup is depicted in the figure below. Now imagine using this set of points to do an ordinary least squares (OLS) regression of the form \\(y = \\beta_0 + x \\beta_1 + \\epsilon\\). What can be said about \\(\\beta_1\\)?\nFigure 1: Rectangle with length \\(2L\\) and width \\(2W\\).\rNote: I don\u0026rsquo;t remember the precise statement of the prompt, but I think this captures the essence of the problem very well.\nSolution\r#\rHere I will address the problem in two ways. The first is an intuitive, qualitative understanding, and the second is a more formal quantitative result which backs up the intuitive solution.\nIntuitive Approach\r#\rOur first guess is that the regressor should give us the straight line that goes through the axis of symmetry of the rectangle parallel to the long side of the rectangle, i.e., the line that makes an angle \\(\\theta\\) with the \\(x\\)-axis and intersects the origin. I will refer to this line as the rectangle\u0026rsquo;s \u0026ldquo;principal axis\u0026rdquo;. It is depicted below for the general case as a cyan line.\nFigure 2: Principal axis illustration.\rIn the specific case \\(\\theta = 0\\) we can be 100% confident the regression line would be the principal axis. However, when \\(\\theta \u0026gt; 0\\) things are not so clear. So let\u0026rsquo;s try to imagine what shape would be guaranteed to give the principal axis as the solution.\nWell, if we take the rectangle with \\(\\theta = 0\\) and perform a skew transformation so that the edges on the left and right stay parallel to the \\(y\\)-axis, then the symmetry properties of the system dictate that the line of best-fit is the principal axis. This is illustrated in the figure below.\nFigure 3: Skewed rectangle and regression line.\rA rectangle that has been rotated about the origin by an angle \\(\\theta\\) can be decomposed into 3 shapes: a skew rectangle and two right-angle triangles as shown below. The skew rectangle \u0026ldquo;votes\u0026rdquo; that the best-fit line is the principal axis as argued earlier. But now there are two more contributions, coming from the right triangles on the left and right of the diagram. If we look at how the principal axis cuts through the rightmost triangle, we observe there is more area below the principal axis, than above, so that triangle \u0026ldquo;votes\u0026rdquo; for the best-fit line to be shifted down. Similarly, for the leftmost triangle there is more area above the principal axis than below, so it \u0026ldquo;votes\u0026rdquo; for the best-fit line to be pulled up on that side. The net effect is a rotation toward a less steep slope.\nFigure 4: Rectangle decomposed into skewed rectangle and two right-triangles.\rSo, we expect the regression slope to be biased toward 0.\nQuantitative Analysis\r#\rI think the intuitive explanation is really satisfying, but this effect seems extremely similar to attenuation bias which occurs when the weak exogeneity assumption of linear regression is violated. I wanted to make this connection more concrete, and this more math-heavy explanation does just that.\nThe entire premise of this problem can be restated in the following way. First let\u0026rsquo;s denote the \u0026ldquo;natural\u0026rdquo; coordinates of the rectangle as \\((u,v)\\). Specifically, \\((u,v) = (0,0)\\) is at the center of the rectangle, \\(u\\) increases along the principal axis of the rectangle, and \\(v\\) is orthogonal to it. More specifically, I\u0026rsquo;ll take it to be in the direction measured \\(\\pi/2\\) radians counter-clockwise from the principal axis. In the figures above \\(v\\) increases as you move up and to the left.\nThe statement that the rectangle has a uniform density of points can be restated as sampling points \\((u,v)\\) uniformly at random from the region \\((u,v) \\in [-L, L] \\times [-W, W]\\). In other words,\n$$ \\begin{align} u \u0026\\sim \\mathrm{U}(-L, L) \\\\ v \u0026\\sim \\mathrm{U}(-W, W) \\ . \\end{align} $$To get their coordinates in terms of \\((x,y)\\) we simply rotate \\((u,v)\\) counter-clockwise by an angle \\(\\theta\\) to yield\n$$ \\begin{align} x \u0026= \\cos \\theta u - \\sin \\theta v \\\\ y \u0026= \\sin \\theta u + \\cos \\theta v \\ . \\end{align} $$The OLS estimator for the slope, \\(\\beta_1\\) is\n$$ \\frac{\\mathrm{Cov}(x,y)}{\\mathrm{Var}(x)} \\ . $$ Since we\u0026rsquo;re regressing on the continuous infinity of points drawn from the distribution, \\(\\mathrm{Var}\\) and \\(\\mathrm{Cov}\\) denote the true variance and covariance. For the covariance we have, $$ \\mathrm{Cov}(x,y) = \\mathbb{E}[x (\\beta_0 + \\beta_1 x + \\epsilon)] - \\mathbb{E}(x) \\mathbb{E}(\\beta_0 + \\beta_1 x + \\epsilon) \\ . $$ Using the expressions we derived for \\(x\\) and \\(y\\) in terms of \\(u\\) and \\(v\\), and noting that \\(\\mathbb{E}(u) = \\mathbb{v} = 0\\), and \\(\\mathrm{Var}(u) = L^2 / 3\\), \\(\\mathrm{Var}(v) = W^2 / 3\\) one can show that $$ \\mathrm{Cov}(x,y) = \\frac{1}{3} (L^2 - W^2) \\cos\\theta \\sin\\theta $$ and $$ \\mathrm{Var}(x) = \\frac{1}{3} (L^2 \\cos^2 \\theta + W^2 \\sin^2\\theta) \\ . $$ Hence, we have $$ \\hat{\\beta}_1 = \\frac{(L^2 - W^2)\\cos\\theta \\sin\\theta}{L^2 \\cos^2\\theta + W^2 \\sin^2\\theta} = \\frac{1 - (W / L)^2}{1 + (W / L)^2 \\tan^2 \\theta} \\tan \\theta \\ . $$ Since we naiively expect the regression line to have a slope near \\(\\tan \\theta\\), the factor in front is a multiplicative bias. Moreover, since here we assume \\(W \u0026lt; L\\), the bias is \\(\u0026lt; 1\\), implying a shallower slope.\nNumerical Verification\r#\rBelow, the cyan-filled rectangle represents the data domain, rotated by an angle \\(\\theta\\) relative to the \\(x\\)-axis. The cyan dotted line shows the principal axis of the rectangle, with a slope of \\(\\tan(\\theta)\\). The red solid line represents the theoretical regression line, calculated based on the derived formula for \\(\\beta_1\\), while the dashed black line corresponds to the OLS regression line obtained from numerical calculations. Each subplot displays the \\(R^2\\) value of the regression, providing a measure of the goodness of fit. The comparison between the theoretical regression slope (red), the numerical regression slope (dashed black), and the principal axis (cyan) highlights the bias introduced by the rectangular geometry. This visualization confirms that the numerical results closely align with the theoretical predictions across varying rectangle dimensions and angles.\nFigure 5: Illustration of the relationship between the OLS regression slope, the principal axis of the rectangle, and the theoretical regression line for datasets confi\r","date":"31 December 2024","externalUrl":null,"permalink":"/posts/a-linear-regression-problem/","section":"Posts","summary":"","title":"A cute linear regression brainteaser","type":"posts"},{"content":" I have started looking into quant interview questions again and I went down a rabbit hole with one I found on quantable.io.\nIt goes like this:\nSuppose you draw points \\(p = (x,y)\\) uniformly at random from the unit square. What is the expected value of the distance \\(s\\) of the sampled point \\(p\\) from the line \\(y = x\\)? The solution to this question is relatively straightforward. First, due to symmetry we can limit ourselves to thinking about one of the triangles formed by dividing the square with the line \\(y = x\\). Next, take the diagonal as the \u0026lsquo;base\u0026rsquo; of the triangle. What we are looking for is the average distance from the base to any point in the triangle. We can express this as an integral, $$ \\bar{y} = \\frac{1}{A} \\int \\int_A y \\mathrm{d}A , $$ where \\(A\\) denotes the area of the triangle. There are many ways to evaluate this integral (you could use Green\u0026rsquo;s theorem, choose coordinate systems that simplify the calculation, etc.) but on my first pass I brute-forced it. First, the area \\(A\\) is just half the area of the unit square, so \\(A = 1/2\\). Thus, $$ \\bar{y} = 2 \\int_0^{1 / \\sqrt{2}} \\int_y^{\\sqrt{2} - y} y \\, \\mathrm{d}x \\mathrm{d}y = \\frac{1}{3\\sqrt{2}} , $$ and we\u0026rsquo;re done.\nAlthough it was relatively straightforward to get the answer, this problem made me ask several more questions. First of all, why is the answer exactly \\(1/3\\) of the height of the triangle? Its simplicity hints that there is an easier way to do the problem. Secondly, the same day that I solved this problem an undergrad asked me to help them with a problem that required using the rotational version of Newton\u0026rsquo;s second law (with torques), and that reminded me that the expectation value we calculated here can be recast as a balancing problem in physics.\nThis was some of the motivation for me to write these notes on centroids.\nWhen I was trying to explain the solution to the student several more things stuck out to me. For one, I wanted to convince them that when we solve problems involving torque, the \\(\\vec{r}\\) we use in \\(\\vec{r} \\times \\vec{F}\\) can be taken to be measured from an arbitrary point, so long as we use the same \\(\\vec{r}\\) to calculate the other torques. Another thought I had was that it would be nice to explain why we define torques (\u0026ldquo;rotational force\u0026rdquo;) as \\(\\vec{r} \\times \\vec{F}\\), and not, say \\(r^{2.31} \\vec{r} \\times \\vec{F}\\). The best answer I could come up with was because of symmetries and Noether\u0026rsquo;s theorem: if we suppose physics is symmetric under \\(SO(3)\\) rotations, we get a corresponding conserved quantity \u0026ndash; angular momentum. Then, we define torque as the thing that causes time-varying changes in angular momentum \u0026ndash; \\(\\vec{\\tau} = \\frac{\\mathrm{d} \\vec{L}}{\\mathrm{d} t}\\). Using \\(\\vec{L} = \\vec{r} \\times \\vec{p}\\), Newton\u0026rsquo;s second law \\(\\vec{F} = \\frac{\\mathrm{d} \\vec{p}}{\\mathrm{d}t}\\), and taking \\(\\vec{r} = \\mathrm{const}\\), we immediately get the usual expression for torque \\(\\vec{\\tau} = \\vec{r} \\times \\vec{F}\\).\n","date":"8 December 2024","externalUrl":null,"permalink":"/posts/quant-problem-and-physics/","section":"Posts","summary":"","title":"A quant probability question and physics","type":"posts"},{"content":"A basic python package featuring one module and some data files.\nproj/ ├─ package/ │ ├─ data/ │ │ ├─ __init__.py │ │ ├─ data_file.csv │ │ ├─ custom_style.mplstyle │ ├─ __init__.py │ ├─ module1.py ├─ MAKEFILE.in ├─ setup.cfg ├─ LICENSE ├─ pyproject.toml ├─ README.md\rMAKEFILE.in\r#\rTells build tool which non-python files to include in the distribution. Note that the file ends in .in, not .ini.\nExample\n# MAKEFILE.in include package/data/*.csv recursive-include package *.mplstyle\rsetup.cfg\r#\rConfiguration options for package builder (e.g. setuptools, poetry, flit.)\nExample (for setuptools)\n[metadata] name = package version = attr: package.__version__ author = John Doe author_email = john.doe@email.com description = A description. long_description = file: README.md long_description_content_type = text/markdown url = https://github.com/username/package classifiers = Programming Language :: Python :: 3 License :: OSI Approved :: MIT License Operating System :: OS Independent [options] packages = find: python_requires = \u0026gt;=3.7 include_package_data = True\rIn this example the version is obtained by importing the package module and reading its .__version__ attribute which must be defined in the __init__.py. In our example it would require something like\n# package/__init__.py __version__ = \u0026#34;0.0.1\u0026#34;\rLICENSE\r#\rVery important. A good resource is https://choosealicense.com/.\npyproject.toml\r#\rExample (for setuptools)\n[build-system] requires = [ \u0026#34;setuptools\u0026gt;=54\u0026#34;, \u0026#34;wheel\u0026#34; ] build-backend = \u0026#34;setuptools.build_meta\u0026#34;\rTo install the package in development mode (so that you don\u0026rsquo;t have to build the package every time you make a change) you can cd to the proj/ directory and run python -m pip install -e ..\nTo distribute the package by uploading it to the python package index (PyPI), first make sure build and twine are installed and make a PyPI account.\npython -m pip install build twine\rpython -m build\rFirst test that your package works on testpypi\ntwine upload -r testpypi dist/*\rThen you can upload it to PyPI with\ntwine upload dist/*\rASCII file trees were made with ASCII Tree Generator.\n","date":"24 July 2022","externalUrl":null,"permalink":"/posts/writing-a-python-package/","section":"Posts","summary":"","title":"Writing a python package","type":"posts"},{"content":"Useful link: https://kb.rice.edu/page.php?id=108237\nConnect to NOTS via SSH.\r#\rIf on campus, or at RGA/RVA, connect to \u0026ldquo;Rice Owls\u0026rdquo; network. (You can\u0026rsquo;t connect to NOTS if you\u0026rsquo;re on \u0026ldquo;Rice Visitor\u0026rdquo;.) If off-campus, follow these instructions to set up the Rice VPN. Then, SSH into NOTS by opening up a terminal and typing\nssh -Y [your NetID, e.g. abc123]@nots.rice.edu\rType in the same password you use to log in to ESTHER. You should now be connected to one of the login nodes.\nOverview of Filestructure\r#\r(screenshot from https://kb.rice.edu/page.php?id=108237)\nAfter SSH-ing into a login node your terminal will be in your local user directory /home/[your NetID]/ e.g. /home/abc123/. You can freely edit anything in this directory.\ncd into $SHARED_SCRATCH and mkdir a directory for yourself:\ncd $SHARED_SCRATCH mkdir [your NetID]\rThe $SHARED_SCRATCH directory is a place to store temporary files for job i/o. Files here may be purged after 14 days. Once a job is done, you will copy the output to your directory in $WORK. First cd into $WORK/[your advisor's NetID] and mkdir [your NetID]. For me I had to write\ncd $WORK/al72 mkdir rmh14\rModules\r#\rmodule avail gives a list of available modules. Adding a keyword afterword, e.g.\nmodule avail math\rreturns a list of modules with the word \u0026ldquo;math\u0026rdquo; in them:\nInteractive session:\r#\rsrun -p=interactive --pty --export=ALL -n=1 -c=1 --time=00:30:00 /bin/bash\rSee https://slurm.schedmd.com/srun.html for explanation of the options. There is a time limit of 30 minutes; if you request 1 hour of time you will receive an error.\nCheck how many resources previous jobs have used with sacct --format=\u0026quot;JobID,CPUTime,MaxRSS\u0026quot;\n","date":"19 November 2021","externalUrl":null,"permalink":"/posts/getting-started-with-rice-nots/","section":"Posts","summary":"","title":"Getting Started with Rice NOTS","type":"posts"},{"content":"\rClick here for solutions PDF\rThis post contains solutions to select problems in Steven Weinberg\u0026rsquo;s \u0026ldquo;The Quantum Theory of Fields: Vol. I\u0026rdquo;. The PDF (link above) was authored by Hong-Yi Zhang, Siyang Ling, Jiazhao Lin, and myself. Please note that this is still a work in progress. If you notice any mistakes please contact me at rmh14 [at] rice [dot] edu.\nCurrently solutions are available for the following problems:\nChapter 2\nProblem 2.2 Problem 2.5 Chapter 3\nProblem 3.3 Problem 3.4 Chapter 4\nProblem 4.1 Problem 4.2 (not finished) Chapter 5\nProblem 5.1 Problem 5.4 Problem 5.5 Chapter 6\nProblem 6.4 Chapter 7\nProblem 7.1 Chapter 9\nProblem 9.1 (Update 13 Feb 2023: solution corrected by Jiangyuan Qian) ","date":"30 August 2021","externalUrl":null,"permalink":"/posts/steven-weinberg-quantum-theory-of-fields-vol-i-solutions-to-selected-problems/","section":"Posts","summary":"Solutions to select problems in Steven Weinberg’s “The Quantum Theory of Fields: Vol. I”.","title":"Weinberg QFT Vol I Solutions","type":"posts"},{"content":"My notes:\nLinear regression Options pricing Triangle wave drive ","externalUrl":null,"permalink":"/notes/","section":"","summary":"","title":"","type":"notes"},{"content":"","externalUrl":null,"permalink":"/authors/","section":"Authors","summary":"","title":"Authors","type":"authors"},{"content":"","externalUrl":null,"permalink":"/categories/","section":"Categories","summary":"","title":"Categories","type":"categories"},{"content":"\rBiography\r#\rAcademic background\r#\rI completed my PhD in Physics at Rice University in Houston. I was advised by Professors Andrew Long and Mustafa Amin. My research consists of quantifying cosmological signatures of axion-like particles, looking for evidence of them in cosmological and astrophysical data. In particular, I focused on the prospect of detecting axion strings via the cosmic microwave background using an effect known as cosmic birefringence. A full list of my publications can be found here, and a copy of my thesis is available here.\nAs an undergrad I did research in computational nanophotonics with Nicolas Large at the University of Texas at San Antonio and did NSF REUs at Brigham Young University and the University of Chicago where I studied quantum dynamics and cosmology respectively.\nA copy of my resume can be found here.\nPersonal background\r#\rMy parents are Japanese and Spanish/American. I grew up in Singapore and moved to the Texas for university, where I\u0026rsquo;ve lived ever since.\nOutside of physics I enjoy origami, photography, and cooking!\n","externalUrl":null,"permalink":"/authors/admin/","section":"Authors","summary":"","title":"Ray Hagimoto","type":"authors"},{"content":"","externalUrl":null,"permalink":"/series/","section":"Series","summary":"","title":"Series","type":"series"},{"content":"","externalUrl":null,"permalink":"/tags/","section":"Tags","summary":"","title":"Tags","type":"tags"}]