<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Prasanth Velithoti's Blog]]></title><description><![CDATA[Prasanth Velithoti's Blog]]></description><link>https://prasanthvelithoti.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Prasanth Velithoti&apos;s Blog</title><link>https://prasanthvelithoti.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 07 Oct 2026 12:27:35 GMT</lastBuildDate><atom:link href="https://prasanthvelithoti.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[9 SQL Interview Questions That Show Whether You've Actually Queried Production Data]]></title><description><![CDATA[The question that ended a candidate's loop last spring wasn't hard. I asked him to find the second-highest salary in an employees table. He wrote ORDER BY salary DESC LIMIT 1 OFFSET 1, and when I aske]]></description><link>https://prasanthvelithoti.hashnode.dev/9-sql-interview-questions-that-show-whether-you-ve-actually-queried-production-data</link><guid isPermaLink="true">https://prasanthvelithoti.hashnode.dev/9-sql-interview-questions-that-show-whether-you-ve-actually-queried-production-data</guid><category><![CDATA[SQL]]></category><category><![CDATA[interview questions]]></category><dc:creator><![CDATA[Prasanth Velithoti]]></dc:creator><pubDate>Wed, 23 Sep 2026 11:22:53 GMT</pubDate><content:encoded><![CDATA[<p>The question that ended a candidate's loop last spring wasn't hard. I asked him to find the second-highest salary in an <code>employees</code> table. He wrote <code>ORDER BY salary DESC LIMIT 1 OFFSET 1</code>, and when I asked what happens if two people share the top salary, he stared at the query for a long ten seconds and said, "It still works?"</p>
<p>It doesn't. It returns the top salary again. And that small gap, the difference between writing SQL that runs and writing SQL that's right, is what most SQL interview rounds are actually testing. Nobody on the other side of the table cares whether you've memorised the syntax for <code>LATERAL</code>. They care whether you think about duplicates, NULLs, and what the query does on the 10-million-row table instead of the 12-row example.</p>
<p>Here are the nine questions I've seen come up most, plus what a strong answer sounds like.</p>
<h2>1. Find the second-highest salary</h2>
<p>The fix for the story above is to stop thinking in rows and start thinking in distinct values:</p>
<pre><code class="language-sql">SELECT MAX(salary)
FROM employees
WHERE salary &lt; (SELECT MAX(salary) FROM employees);
</code></pre>
<p>Or, more generally, with a window function:</p>
<pre><code class="language-sql">SELECT salary
FROM (
  SELECT salary, DENSE_RANK() OVER (ORDER BY salary DESC) AS rnk
  FROM employees
) t
WHERE rnk = 2
LIMIT 1;
</code></pre>
<p>What the interviewer wants to hear is <em>why</em> you picked <code>DENSE_RANK</code> and not <code>ROW_NUMBER</code>. Then say, unprompted, what happens if there's no second salary: the first query returns NULL, the second returns no rows. Mentioning that tells them you've been burned by an empty result set feeding a dashboard before.</p>
<h2>2. <code>ROW_NUMBER</code> vs <code>RANK</code> vs <code>DENSE_RANK</code></h2>
<p>The weak answer lists definitions. The strong answer uses one dataset to show all three. Salaries of 100, 100, 90 give:</p>
<ul>
<li><code>ROW_NUMBER</code>: 1, 2, 3 (arbitrary tie-break, so it's not deterministic unless you add a tiebreaker column)</li>
<li><code>RANK</code>: 1, 1, 3 (it skips)</li>
<li><code>DENSE_RANK</code>: 1, 1, 2 (no gaps)</li>
</ul>
<p>The follow-up almost always goes: "When would you use <code>ROW_NUMBER</code>?" The answer that lands is deduplication. You keep the latest record per customer with <code>ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY updated_at DESC) = 1</code>. That's a real, everyday use, and it shows you've cleaned real data.</p>
<h2>3. <code>WHERE</code> vs <code>HAVING</code></h2>
<p>Everyone knows <code>HAVING</code> filters after grouping. Fewer people can explain the order of logical evaluation: <code>FROM</code> → <code>WHERE</code> → <code>GROUP BY</code> → <code>HAVING</code> → <code>SELECT</code> → <code>ORDER BY</code> → <code>LIMIT</code>. That order is also why you can't use a <code>SELECT</code> alias inside <code>WHERE</code> in most databases.</p>
<p>Here's the practical point that separates people: push every filter you can into <code>WHERE</code>. Filtering 5 million rows down to 50,000 <em>before</em> aggregation is cheaper than aggregating everything and throwing groups away in <code>HAVING</code>. If you say that, you've shown you think about cost, not just correctness.</p>
<h2>4. Explain the <code>JOIN</code> types, and the NULL trap in <code>LEFT JOIN</code></h2>
<p>Inner, left, right and full outer are table stakes. The question is really a setup for this bug:</p>
<pre><code class="language-sql">SELECT c.id, o.id
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id
WHERE o.status = 'paid';
</code></pre>
<p>That <code>WHERE</code> clause silently turns the left join into an inner join, because customers with no orders have <code>o.status = NULL</code>, and <code>NULL = 'paid'</code> isn't true. If you meant "all customers, plus their paid orders", the condition belongs in the <code>ON</code> clause. I've seen this exact bug under-report a churn metric for a month in production. Candidates who spot it without a hint almost always get a strong signal.</p>
<h2>5. How does <code>NULL</code> behave in comparisons and aggregates?</h2>
<p>Short version: <code>NULL</code> means unknown, so <code>NULL = NULL</code> is not true, and you need <code>IS NULL</code>. <code>COUNT(*)</code> counts rows, but <code>COUNT(column)</code> skips NULLs. <code>AVG</code> ignores NULLs too, which means the average of 10, NULL, 20 is 15, not 10.</p>
<p>The trap to mention is <code>NOT IN</code> with a subquery that contains a NULL. It returns no rows at all. <code>NOT EXISTS</code> doesn't have that problem, which is one reason many teams default to it. This is the kind of answer that shows you've debugged a query that "should" return data and didn't.</p>
<h2>6. What's an index, and when does it <em>not</em> help?</h2>
<p>Most candidates can say "a B-tree that speeds up lookups." The interesting half of the answer is when an index gets ignored:</p>
<ul>
<li>A function on the column: <code>WHERE LOWER(email) = ...</code> won't use a plain index on <code>email</code> (you'd need an expression index)</li>
<li>A leading wildcard: <code>LIKE '%gmail.com'</code></li>
<li>Low-selectivity columns, where scanning the table is cheaper anyway</li>
<li>Composite indexes queried without the leftmost column: an index on <code>(country, city)</code> helps <code>WHERE country = 'IN'</code> but not <code>WHERE city = 'Pune'</code></li>
</ul>
<p>Then add the cost side: every index slows down writes and takes storage. My one mild disagreement with common advice goes here. "Index every column you filter on" is bad advice for write-heavy tables. I've seen an insert-heavy events table get noticeably faster after we <em>dropped</em> four indexes nobody was reading from.</p>
<h2>7. Read this query plan</h2>
<p>More interviews now paste in an <code>EXPLAIN ANALYZE</code> output and ask what's wrong. You don't need to know every node type. Look for three things:</p>
<ol>
<li>A <code>Seq Scan</code> on a big table where you expected an index scan</li>
<li>A big gap between estimated rows and actual rows, which usually means stale statistics (run <code>ANALYZE</code>)</li>
<li>A nested loop joining two large sets, where a hash join would be expected</li>
</ol>
<p>Just say "I'd check whether the row estimates match the actuals first." That tells them you've read real plans, not just tutorials.</p>
<h2>8. Transactions and isolation levels</h2>
<p>Explain ACID briefly, then get concrete about isolation, because that's where the real questions live. Name the anomalies (dirty reads, non-repeatable reads, phantom reads) and which isolation level prevents which. Be ready for the PostgreSQL detail: its default is Read Committed, and its Repeatable Read is implemented with snapshots, so it also prevents phantoms in practice.</p>
<p>The follow-up is usually a scenario. Two requests withdraw from the same account at the same time: how do you stop the balance going negative? Good answers are <code>SELECT ... FOR UPDATE</code>, a conditional update like <code>UPDATE accounts SET balance = balance - 50 WHERE id = 1 AND balance &gt;= 50</code> (then check rows affected), or serializable isolation with a retry. Offering the conditional update first shows pragmatism.</p>
<h2>9. Write a running total and a 7-day moving average</h2>
<p>This one is the window-function capstone:</p>
<pre><code class="language-sql">SELECT
  order_date,
  daily_revenue,
  SUM(daily_revenue) OVER (ORDER BY order_date) AS running_total,
  AVG(daily_revenue) OVER (
    ORDER BY order_date
    ROWS BETWEEN 6 PRECEDING AND CURRENT ROW
  ) AS moving_avg_7d
FROM daily_sales;
</code></pre>
<p>The detail that impresses is the gap problem. <code>ROWS BETWEEN 6 PRECEDING</code> means seven <em>rows</em>, not seven <em>days</em>. If there are missing dates, your "7-day" average quietly covers ten days. Mention generating a date series and left-joining onto it (or using <code>RANGE</code> with an interval where your database supports it). Almost nobody brings this up without being asked.</p>
<h2>What actually goes wrong in the live round</h2>
<p>The candidates I've watched struggle usually know this material. What happens is that they freeze when they're writing SQL in a shared editor while someone watches every keystroke, and the NULL edge case they'd catch at their desk just doesn't come to mind. Two things help. First, narrate your assumptions out loud ("I'm assuming salary can't be NULL. Should I handle that?"). Second, practise under the same pressure you'll face on the day. Some people now also keep a real-time assistant such as <a href="https://lastroundai.com/products/ai-interview-copilot">LastRound AI's interview copilot</a> running as a safety net for the moments when a detail like <code>DENSE_RANK</code> vs <code>RANK</code> slips mid-sentence. Whether you use one or not, the fix is the same: rehearse the edge cases until you mention them without being asked.</p>
<p>One honest caveat. Everything above leans toward PostgreSQL behaviour. MySQL, SQL Server and Snowflake differ in the details (default isolation levels, <code>LIMIT</code> vs <code>TOP</code>, how NULLs sort), so ask which engine the team uses before you commit to syntax. That question alone scores points.</p>
]]></content:encoded></item></channel></rss>