Tai-Ning Liao
Home
Blog
Projects
ARC
Categories
All
(4)
ReinforcementLearning
(3)
Blog
我被PPO卡住了
原本以為 PPO (proximal policy optimization) 最大的貢獻是 important sampling。但那大家早就會了。
Aug 14, 2026
如果要做一台會思考的機器
現在的機器是怎麼思考的?
Aug 11, 2026
Policy, Q 和 V:到底該用哪個
【師】「今天要講的大概是 RL,Day 1 到 Day 10 的部分。我現在大概整理出一個脈絡,這個脈絡既不是照歷史的順序,也不是我心中最完美的走法。我心中比較完美的走法是這樣:RL 是一個很大的…
Aug 6, 2026
Hello, world
Aug 5, 2026
No matching items