聖塔非研究所

摘要 The practical successes of deep 神經 網絡s have not b

2019-12-01 · 已發表論文 · 更新 2026/08/30 下午12:48

摘要 The practical successes of deep 神經 網絡s have not been matched by theoretical progress that satisfyingly explains their behavior. In this work, we study the 資訊 bottleneck (IB) theory of 深度學…

本頁只刊出中文翻譯與中文說明;英文原文請見下方原文連結。

原文連結

論文資訊

  • 類型:已發表論文
  • 日期:2019-12-01

摘要

The practical successes of deep 神經 網絡s have not been matched by theoretical progress that satisfyingly explains their behavior. In this work, we study the 資訊 bottleneck (IB) theory of 深度學習, which makes three specific claims: first, that deep 網絡s undergo two distinct phases consisting of an initial fitting phase and a subsequent compression phase; second, that the compression phase is causally related to the excellent generalization performance of deep 網絡s; and third, that the compression phase occurs due to the 擴散-like behavior of 隨機 gradient descent. Here we show that none of these claims hold true in the general case, and instead reflect assumptions made to compute a finite mutual 資訊 metric in deterministic 網絡s. When computed using simple binning, we demonstrate through a combination of

※ 此為已發表論文,全文需透過期刊付費取得