<feed xmlns="http://www.w3.org/2005/Atom"> <id>https://roba269.github.io/</id><title>Liang He</title><subtitle>Liang's learning notes</subtitle> <updated>2026-04-03T23:58:28-07:00</updated> <author> <name>Liang He</name> <uri>https://roba269.github.io/</uri> </author><link rel="self" type="application/atom+xml" href="https://roba269.github.io/feed.xml"/><link rel="alternate" type="text/html" hreflang="en" href="https://roba269.github.io/"/> <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator> <rights> © 2026 Liang He </rights> <icon>/assets/img/favicons/favicon.ico</icon> <logo>/assets/img/favicons/favicon-96x96.png</logo> <entry><title>nano-vllm code walkthrough</title><link href="https://roba269.github.io/posts/nano-vllm-code-walkthrough/" rel="alternate" type="text/html" title="nano-vllm code walkthrough" /><published>2026-04-03T00:00:00-07:00</published> <updated>2026-04-03T23:58:07-07:00</updated> <id>https://roba269.github.io/posts/nano-vllm-code-walkthrough/</id> <content type="text/html" src="https://roba269.github.io/posts/nano-vllm-code-walkthrough/" /> <author> <name>Liang He</name> </author> <category term="Inference" /> <summary>Code study of nano-vllm, a minimal implementation of vLLM</summary> </entry> <entry><title>A study on CUDA async memcpy</title><link href="https://roba269.github.io/posts/a-study-on-cuda-async-memcpy/" rel="alternate" type="text/html" title="A study on CUDA async memcpy" /><published>2026-02-27T00:00:00-08:00</published> <updated>2026-03-01T15:46:53-08:00</updated> <id>https://roba269.github.io/posts/a-study-on-cuda-async-memcpy/</id> <content type="text/html" src="https://roba269.github.io/posts/a-study-on-cuda-async-memcpy/" /> <author> <name>Liang He</name> </author> <category term="CUDA" /> <summary>A study on CUDA async exeuctions, including PTX and C++ barrier/pipeline abstractions</summary> </entry> </feed>
