<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Pytorch on Reza Ghafari</title>
    <link>https://rezag.io/tags/pytorch/</link>
    <description>Recent content in Pytorch on Reza Ghafari</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Fri, 29 May 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://rezag.io/tags/pytorch/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Self-Attention from Scratch: The Heart of the Transformer</title>
      <link>https://rezag.io/posts/2026/self-attention-from-scratch/</link>
      <pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate>
      <guid>https://rezag.io/posts/2026/self-attention-from-scratch/</guid>
      <description>I&amp;rsquo;m writing this post mostly to make sure this stuff actually sticks in my head. Step-by-step explanations of self-attention with working examples are hard to find.&#xA;If you have used ChatGPT, Claude, or any modern language model, you have interacted with a Transformer. The Transformer is the architecture behind virtually every large language model (LLM) today. It was introduced in the legendary 2017 paper &amp;ldquo;Attention Is All You Need&amp;rdquo; by Vaswani et al.</description>
    </item>
  </channel>
</rss>
