IT / 文献库

PAPER 165 / CLOSE READING

Invariant visual object recognition: biologically plausible approaches

语义审核:pass · 图表审核:pass

本文目录 (Table of Contents)
  1. 研究背景
  2. 研究思路
  3. 方法
  4. 主要结果
  5. 图注解读
    1. 图 1 · 会聚架构:从脑到 VisNet
    2. 图 2 · HMAX 的 S-C 层级示意
    3. 图 3 · Caltech-256 例图
    4. 图 4 · 实验一分类成绩柱状图
    5. 图 5 · 单神经元放电率的三方对照
    6. 图 6 · ALOI 例图
    7. 图 7 · 实验二分类成绩曲线
    8. 图 8 · 训练集上的单神经元响应
    9. 图 9 · 测试集上的泛化
    10. 图 10 · 相似度矩阵:表征几何的终审
    11. 图 11 · 打乱面孔刺激
    12. 图 12 · 打乱前后的神经元响应对照
    13. 图 13 · 两只杯子的四个视角
    14. 图 14 · 杯子任务的单神经元终局对比
  6. 讨论
  7. 一句话总结
  8. 审校与证据追溯 (Verification & Evidence)
    1. 图表审计结果
    2. 关键事实与局限性声明

这是一篇"用同一套考题给两大模型打分"的建模对比研究:作者把腹侧视觉通路不变物体识别的两个代表性方案——Rolls 系的 VisNet(含时间痕迹学习规则)与 Poggio 系的 HMAX(S-C 交替层级)——放进四个实验里,逐条对照下颞叶(IT)皮层神经元的五项关键性质。结果是 VisNet 在不变表征、稀疏性、整物体构型编码与局部学习规则上全面贴近 IT 神经元,而 HMAX 的输出稀疏性差、几乎不含单神经元信息,识别必须外包给生物学上说不通的 SVM,因为它没有任何学习视角不变性的机制。

研究背景

灵长类视觉系统的一项核心成就是构建对尺寸、对比、视网膜位置、视角、光照等变换相对独立的物体表征——即不变表征(invariant representation)。作者强调这不只是知觉问题:只有 IT 提供了不变表征,下游脑区才可能一次学习就掌握某个物体与奖惩、位置、新旧经验的联结并推广到它的其他视角。而计算模型正是检验"皮层如何完成这一计算"的基本手段。

作者先列出 IT 神经元需要被模型满足的五项关键性质:(1)变换不变性(transform invariance,涵盖平移、尺寸、对比、旋转乃至视角);(2)稀疏分布式表征(sparse distributed representation),单个神经元只对少数刺激高放电,单神经元放电率即可读出较多信息,且至少到几十个神经元的规模上各自独立编码;(3)许多 IT 神经元对"整物体"响应,对单独呈现的部件或打乱空间构型的部件不响应;(4)编码个体物体/面孔的身份,而不仅是"脸 vs 非脸"这类类别——这一点恰是许多模型(包括当时热门的深度网络工作)回避的;(5)学习机制必须是局部的突触学习规则,HMAX、新感知机(neocognitron)与卷积网络中的"权重横向复制"因非局部而不可信。缺口在于:VisNet 与 HMAX 这两类主流方案从未被放在同一组实验里按这五条逐一打分。

研究思路

作者的逻辑是"以性质为纲、以实验为考题":不比谁在基准数据集上分数高,而是问每个模型的神经元输出是否像真正的 IT 神经元——是否对同一物体的全部变换保持相似放电、是否稀疏、是否编码整物体构型与个体身份、能否被生物学上合理的读出机制(如模式联想网络 pattern association network 的点积解码)读取。四个实验层层递进:先用 Caltech-256 这类"随机同类别范例"基准检验分类与表征稀疏性;再换 ALOI 这类"同一物体系统性变换视角"的数据集检验视角不变学习;然后用打乱面孔部件检验是否真的编码形状构型;最后用"视角变换时图像发生灾变式改变"的杯子检验模型能否跨越最难的变换。

选择 VisNet 与 HMAX 作对照组有其设计意图:两者都以 Gabor 滤波近似 V1 输入、都声称"生物学启发",但机制迥异——VisNet 用竞争学习加痕迹学习规则(trace learning rule)在层级内部自组织出不变表征,而 HMAX 的 S 层做模板匹配、C 层做 MAX 池化,不变性全靠层级末端的强力分类器补足。这个对比因此不是"谁分数高",而是"不变性到底长在层级里还是长在解码器里"。

方法

VisNet 为四层层级竞争网络(每层 128×128,共 65,536 个神经元;第 1–4 层分别对应 V2、V4、后部 IT/TEO、前部 IT/TE),输入为 16 种 Gabor 滤波器组(空间频率 0.0625–0.5 周/像素、四个朝向、正负相位),前传连接按高斯概率局部收敛使感受野逐层扩大(第 1 层 100 个连接,第 2–4 层各 400 个)。每层内做横向抑制与 sigmoid 对比增强,用阈值把群体稀疏性控制在设定值;学习采用改进的 Hebb 型痕迹规则 δw = α·y(τ−1)·x(τ)——即用前一时刻活动的衰减痕迹与当前输入结合,其生物学基础被归于刺激呈现后持续 100–400 ms 的持续放电、NMDA 受体约 100 ms 的谷氨酸结合窗口及一氧化氮等弥散信号;痕迹参数 η 第 2 层 0.6、第 3/4 层 0.8,在方法段注明"一般训练 50 epoch;实验 3 为 20 epoch、实验 4 为 10 epoch"。与旧版相比,本文的 VisNet 用了 Gabor 输入、更大的网络规模,并引入模式联想解码(每类一个输出神经元、每类取 10 个最选择性输入神经元、Hebb 规则训练一次)作为生物学上合理的读出方式,配合与 IT 电生理完全同款的香农信息量指标(单细胞刺激特异性信息 I(s,R) 与多细胞互信息)。

HMAX 采用 Mutch & Lowe (2008) 的实现(两层 S-C;补充材料用 Serre et al. 2007a 的三层版本复核,结论一致),其 S 单元把突触权重设为随机抽取的训练图像小块(模板匹配)、C 单元对同类 S 单元取 MAX,且滤波器在层内横向复制;标准做法是在 C2 层输出上接支持向量机(SVM)。为模拟 Serre 等人声称对应后部 IT 的视调单元(view-tuned unit, VTU)层,作者按 Riesenhuber & Poggio (1999) 的方式为每个训练视角/范例设立一个 VTU,再经一层最小二乘感知机分类。作者特别提醒:HMAX 家族约有一千万个计算单元,至少是 VisNet(65,536)的一百倍,这一点在解读性能对比时需要记住。

主要结果

  1. 实验 1(Caltech-256 分类):分数相当,表征天差地别(图 3–5)。对帽子 vs 啤酒杯两类做 SVM 分类,VisNet 与 HMAX C2 成绩相似且都不高(两类内图像差异太大,交叉验证困难)。但表征层面分道扬镳:VisNet 输出神经元稀疏——一个神经元对 10 个未训练帽子范例中的 8 个高放电、对啤酒杯只响应 1 个(单细胞信息 0.38 bit,五细胞均值 0.28 bit);HMAX 的 C2 神经元对两个类的几乎每张图都同样高放电(归一化均值 0.905 vs 0.900,五细胞均值信息仅 0.07 bit),VTU 层同样不理想(均值 0.10 bit)。换用生物学合理的模式联想解码后,两者只有 61.7%–63% 的成绩(随机水平 50%),且 HMAX 的"成绩"主要靠外部强分类器挣来。作者由此得出方法论结论:Caltech 型随机范例数据集根本不提供学习视角不变性所需的信息。

  2. 实验 2(ALOI 视角不变学习):VisNet 学得会,HMAX 学不会(图 6–10)。8 个物体、每个 72 张每 5° 一张的转台图像;训练集为每物体 4 个视角(间隔 90°)、9 个(40°)或 18 个(20°),测试集为与之错开 10° 的 18 个视角(随机水平 12.5%)。VisNet 只要有少量训练视角就大幅越过随机(未训练网络仅 18%,训练 9 个视角后 73%),且在离最近训练视角 20° 的中间视角上也能保持约 72–77% 的成绩;最佳单神经元对物体 4(灯泡)的全部 9 个训练视角都响应、对其他物体全不响应,单细胞信息达 3 bit(五细胞均值 2.2 bit)。HMAX 的 C2 单元即便对每类最"挑"的那个单元也只对单一视角有反应(0.68 bit,均值 0.28 bit);补一句:"(18 个训练视角时 VisNet 用 SVM 同样达 92%,再次说明差异在表征而非读出器)"。相似度矩阵(图 10)更直观:VisNet 的输出在同一物体的各视角间高度相似、在不同物体间接近 0;HMAX C2 的所有输出之间相关性最低也有 0.975——同一物体的不同视角与其他物体一样"远"。

  3. 实验 3(打乱面孔部件):VisNet 编码的是构型,HMAX 靠纹理(图 11、12)。用 8 张 ORL 面孔、每张 5 个视角训练,再把每张面孔的四个象限随机重排。VisNet 训练集 100% 正确,但测试打乱面孔时降到随机水平(12.5%,多细胞信息 0.0 bit),示例神经元的放电率直接归零——与 IT 神经元对构型破坏的敏感性一致。HMAX 的视调神经元对打乱版维持与原图同样高的放电,说明其判别依据是残存的低层特征与纹理而非形状;作者解释这 Riesenhuber & Poggio (1999) 当年用回形针类简化图形报告的打乱效应下降并不矛盾,因为对自然图像打乱恰恰保留了纹理信息。

  4. 实验 4(灾变性视角变换):连续性假设的加分题(图 13、14)。两只 Blender 渲染的杯子(各 4 个视角),有的视角能看到杯座上的文字、有的看不到——图像统计性质在变换中剧变。VisNet 训练 10 个 epoch 后 100% 正确,第 4 层出现对"Bill"全部视角响应、对"Jane"全不响应(或反之)的神经元;HMAX 的 C2 与 VTU 层成绩都是 50%(即随机)、0.0 bit,其神经元放电由"图中有没有文字"主导。这说明 HMAX 的 S-C 层级没有学习"哪些图像属于同一物体"的机制,输出反映的是单张图像的统计特性。

图注解读

图 1 · 会聚架构:从脑到 VisNet

原文图注:Fig. 1 Convergence in the visual system. Right as it occurs in the brain. V1, visual cortex area V1; TEO, posterior inferior temporal cortex; TE, inferior temporal cortex (IT). Left as implemented in VisNet. Convergence through the network is designed to provide fourth layer neurons with information from across the entire input retina

图中左右并列:右侧是真实腹通路各层(V1、V2、V4、TEO、TE)感受野随偏心距增大的实测尺度(约 1.3°→50°),左侧是 VisNet 四层的对应设计——通过逐层会聚让第 4 层神经元得以接收来自整个输入网膜的信息。这张图给出全文的解剖学坐标系:VisNet 的层级不是任意的,而是逐层映射 V2→V4→TEO→TE 的会聚结构。

Figure 1

图 2 · HMAX 的 S-C 层级示意

原文图注:Fig. 2 Sketch of Riesenhuber and Poggio (1999) HMAX model of invariant object recognition. The model includes layers of 'S' cells, which perform template matching (solid lines), and 'C' cells (solid lines), which pool information by a non-linear MAX function to achieve invariance (see text) (After Riesenhuber and Poggio 1999.)

这是对照模型的说明书:S 细胞层做模板匹配(把突触权重设为随机范例的输入),C 细胞层用非线性 MAX 函数池化以获得一定的不变性,S-C 交替铺成层级,末端通常接 SVM 分类。读懂它才能理解后文对 HMAX 的两项指控——模板匹配 + 权重横向复制(非局部操作)+ 强分类器兜底,正是其输出"不像 IT"的根源。

Figure 2

图 3 · Caltech-256 例图

原文图注:Fig. 3 Example images from the Caltech256 database for two object classes, hats and beer mugs

帽子与啤酒杯两类各几张范例的缩略展示:同一类内姿态、光照、尺度、遮挡差异巨大。它直观呈现实验 1 用的"随机同类别范例"式数据集长什么样,也解释了为何 SVM 成绩都不高——这类数据没有提供同一物体连续变换的信息。

Figure 3

图 4 · 实验一分类成绩柱状图

原文图注:Fig. 4 Performance of HMAX and VisNet on the classification task (measured by the proportion of images classified correctly) using the Caltech-256 dataset and linear support vector machine (SVM) classification. The error bars show the standard deviation of the means over three cross-validation trials with different images chosen at random for the training set on each trial. There were two object classes, hats and beer mugs, with the number of training exemplars shown on the abscissa. There were 30 test examples of each object class. All cells in the C2 layer of HMAX and layer 4 of Visnet were used to measure the performance. Chance performance at 50% is indicated

横轴为每类训练范例数(5/15/30),纵轴为 SVM 分类正确率,VisNet 与 HMAX C2 两条曲线交织、均不亮眼,50% 为随机水平。此图传达实验 1 的第一个信息:在基准数据集的"记分"层面两个模型差不多——真正的分野要在下一张图的"神经元像不像"层面才显现。

Figure 4

图 5 · 单神经元放电率的三方对照

原文图注:Fig. 5 Top firing rate of two output layer neurons of VisNet, when tested on two of the classes, hats and beer mugs, from the Caltech 256. The firing rates to 10 untrained (i.e. testing) exemplars of each of the two classes are shown. One of the neurons responded more to hats than to beer mugs (solid line). The other neuron responded more to beer mugs than to hats (dashed line). Middle firing rate of two C2 tuned units of HMAX when tested on two of the classes, beer mugs and hats, from the Caltech 256. Bottom firing rate of a view-tuned unit of HMAX when tested on two of the classes, hats (solid line) and beer mugs (dashed line), from the Caltech 256. The neurons chosen were those with the highest single cell information that could be decoded from the responses of a neuron to 10 exemplars of each of the two objects (as well as a high firing rate) in the cross-validation design

三栏都是"归一化放电率 vs 范例编号"的折线图,选取的是各类信息量最高的神经元:上(VisNet)两条曲线在两类之间明显分离;中(HMAX C2)两条曲线对两类几乎无差别地一起高位运行——这就是"非稀疏、无单细胞信息"的可视化;下(HMAX VTU)分离也不清晰。支撑主结果第 1 条中表征质量的核心判据。

Figure 5

图 6 · ALOI 例图

原文图注:Fig. 6 Example images from the two object classes within the ALOI database, a 293 (light bulb) and b 156 (clock). Only the 45◦increments are shown

灯泡(293)与时钟(156)在转台上每 45° 一张的示例。与图 3 恰成对照:ALOI 提供的是"同一物体"的系统视角变换,相邻图像间有连续性——这正是 VisNet 痕迹规则赖以工作的输入统计,也是实验 2 的设计动机。

Figure 6

图 7 · 实验二分类成绩曲线

原文图注:Fig. 7 Performance of VisNet and HMAX C2 units measured by the percentage of images classified correctly on the classification task with 8 objects using the Amsterdam Library of Images dataset and measurement of performance using a pattern association network with one output neuron for each class. The training set consisted of 4 views of each object spaced 90◦apart; or 9 views spaced 40◦apart; or 18 views spaced 20◦apart. The test set of images was in all cases a cross-validation set of 18 views of each object spaced 20◦apart and offset by 10◦from the training set with 18 views and not including any training view. The 10 best cells from each class were used to measure the performance. Chance performance was 12.5% correct

横轴为每类训练视角数(0–19),纵轴为 8 类分类正确率(随机 12.5%),VisNet 与 HMAX C2 各一条曲线,读出均为模式联想网络。VisNet 曲线在极少训练视角时就大幅跃起并稳定在 70% 上下,HMAX C2 则长期贴近随机。这是主结果第 2 条的成绩版证据。

Figure 7

图 8 · 训练集上的单神经元响应

原文图注:Fig. 8 Top firing rate of one output layer neuron of VisNet, when trained on 8 objects from the Amsterdam Library of Images, with 9 views of each object spaced 40◦apart. The firing rates on the training set are shown. The neuron responded to all 9 views of object 4 (a light bulb) and to no views of any other object. The neuron illustrated was chosen to have the highest single cell stimulus-specific information about object 4 that could be decoded from the responses of the neurons to all 72 exemplars shown, as well as a high firing rate to object 4. Middle firing rate of one C2 unit of HMAX when trained on the same set of images. The unit illustrated was that the highest mean firing rate across views to object 4 relative to the firing rates across all stimuli and views. Bottom firing rate of one view-tuned unit (VTU) of HMAX when trained on the same set of images. The unit illustrated was that the highest firing rate to one view of object 4

三栏柱状/点线图展示训练集上"每类最挑"神经元的放电率(横轴为 8 个物体、每物体 9 个视角):上栏 VisNet 神经元对灯泡全部 9 个视角整齐高放、对其他物体全零(信息 3 bit);中栏 C2 单元只在个别视角有一点反应(0.68 bit);下栏 VTU 也只认准一个视角。不变表征"长在神经元里"与"长在解码器里"的差别在此一目了然。

Figure 8

图 9 · 测试集上的泛化

原文图注:Fig. 9 Top firing rate during cross-validation testing of one output layer neuron of VisNet, when trained on 8 objects from the Amsterdam Library of Images, with 9 exemplars of each object with views spaced 40◦apart. The firing rates on the cross-validation testing set are shown. The neuron was selected to respond to all views of object 4 of the training set, and as shown responded to 7 views of object 4 in the test set each of which was 20◦from the nearest training view and to no views of any other object. Middle firing rate of one C2 unit of HMAX when tested on the same set of images. The neuron illustrated was that the highest mean firing rate across training views to object 4 relative to the firing rates across all stimuli and views. The test images were 20◦away from the test images. Bottom firing rate of one view-tuned unit (VTU) of HMAX when tested on the same set of images. The neuron illustrated was that the highest firing rate to one view of object 4 during training. It can be seen that the neuron responded with a rate of 0.8 to the two training images (1 and 9) of object 4 that were 20◦away from the image for which the VTU had been selected

与图 8 同款布局,但考的是训练中从未出现、与最近训练视角相差 20° 的中间视角:VisNet 神经元对灯泡 9 个测试视角中的 7 个响应、其余物体全零——泛化成功;C2 单元对测试视角基本失灵;VTU 只对恰好紧邻其训练视角的两张图有反应。此图证明 VisNet 的不变性是学到的表征而非对训练图的记忆,支撑主结果第 2 条。

Figure 9

图 10 · 相似度矩阵:表征几何的终审

原文图注:Fig. 10 Similarity between the outputs of the networks between the 9 different views of 8 objects produced by VisNet (top), HMAX C2 (middle), and HMAX VTUs (bottom) for the Amsterdam Library of Images test. Each panel shows a similarity matrix (based on the cosine of the angle between the vectors of firing rates produced by each object) between the 8 stimuli for all output neurons of each type. The maximum similarity is 1, and the minimal similarity is 0

三张 72×72 余弦相似度矩阵(行为刺激、列为网络输出向量):VisNet 呈现清晰的块状对角结构——同一物体各视角间高度相似、跨物体接近 0;HMAX C2 整片通红(最低相关 0.975),说明其输出对什么都"像";VTU 介于其间但类内类外难分。这是主结果第 2 条中"HMAX 需要外部强分类器"论断的最直接证据。

Figure 10

图 11 · 打乱面孔刺激

原文图注:Fig. 11 Examples of images used in the scrambled faces experiment. Top two of the 8 faces in 2 of the 5 views of each. Bottom examples of the scrambled versions of the faces

上排为 8 张 ORL 面孔中的两张及其视角范例,下排为把四个象限随机重排后的打乱版。它定义了实验 3 的自变量:部件俱在、低层特征与纹理不变,唯独空间构型被破坏——因此任何对打乱版仍强响应的模型,都可以被判定为没有编码形状。

Figure 11

图 12 · 打乱前后的神经元响应对照

原文图注:Fig. 12 Top effect of scrambling on the responses of a neuron in VisNet. This VisNet layer 4 neuron responded to one of the faces after training and to none of the other 7 faces. The neuron responded to all the different view exemplars 1–5 of the unscrambled face (exemplar normal). When the same neuron was then tested with the randomly scrambled versions of the same face stimuli (exemplar scrambled), the firing rate was zero. Bottom effect of scrambling on the responses of a neuron in HMAX. This view-tuned neuron of HMAX was chosen to be as discriminating between the 8 face identities as possible. The neuron responded to all the different view exemplars 1–5 of the unscrambled face. When the same neuron was then tested with the randomly scrambled versions of the same face stimuli, the neuron responded with similarly high rates to the scrambled stimuli

两栏折线图各分"正常/打乱"两段:上(VisNet)正常段对目标面孔 5 个视角高放、打乱段放电归零;下(HMAX)打乱段维持同样高的放电。效应方向完全相反,直接支撑主结果第 3 条——VisNet 学的是构型敏感的整物体形状,HMAX 依赖的是低层特征与纹理。

Figure 12

图 13 · 两只杯子的四个视角

原文图注:Fig. 13 View-invariant representations of cups. The two objects, each with four views

"Bill"与"Jane"两只杯子各四个视角的渲染图,部分视角能看到杯座文字、部分看不到。这组刺激制造了"灾变性视角变换":同一物体在不同视角下的图像统计可以截然不同,专为检验模型能否跨越图像相似性直接绑定物体身份,是实验 4 的自变量说明。

Figure 13

图 14 · 杯子任务的单神经元终局对比

原文图注:Fig. 14 Top view-invariant representations of cups. Single cells in the output representation of VisNet. The two neurons illustrated responded either to all views of one cup (labelled 'Bill') and to no views of the other cup (labelled 'Jane'), or vice versa. Middle single cells in the C2 representation of HMAX. Bottom single cells in the view-tuned unit output representation of HMAX

三栏折线(横轴为 Bill 与 Jane 各 4 个范例):上栏 VisNet 出现完美的"全视角选一物"神经元;中下栏 HMAX 的 C2 与 VTU 神经元对两只杯子的响应被"图中是否有文字"支配、对物体身份无辨别力。配合正文 50% 正确、0.0 bit 的群体读数,支撑主结果第 4 条。

Figure 14

讨论

作者在讨论中按开篇列出的五项 IT 性质逐条结算:变换不变性、稀疏分布式编码、构型敏感的整物体响应、个体身份编码与局部学习规则,VisNet 全部达标,而 HMAX 在每一条上都需要外挂——要么靠 SVM 学习分类,要么靠每个训练视角一个 VTU 再加一层最小二乘感知机,这些操作(非局部权重复制、有教师的最小二乘学习)在新皮层中都没有对应物。作者由此把批判的矛头温和地指向更广的模型家族:卷积网络的权重共享同样非局部;以表征相似性结构(Khaligh-Razavi and Kriegeskorte 2014;Cadieu et al. 2014)评价模型的做法,回避了"IT 神经元编码个体身份"这一根本维度,且其成绩往往得益于监督训练。

VisNet 路线的理论内核是把时空连续性当作免费教师:灵长类注视一个物体的短时间内,该物体以平移、缩放、旋转、视角变化的形式连续呈现,痕迹学习规则据此把"时间上相继出现的图像"绑定为同一物体的不同变换。作者援引了闭合这一推理的实验链条——把 3D 物体放进猕猴生活环境、无奖赏条件下即可产生视角不变神经元(Booth and Rolls 1998),时间连续性假说亦被 Li & DiCarlo (2008, 2010, 2012) 与人类心理物理学(Perry et al. 2006;Wallis 2013)证实。局限方面文中承认:V1 输入不经训练、Gabor 滤波器组固定、输入图像统一缩放。或直接删去"尺度固定",且比较对象 HMAX 家族参数由原作者代码定夺、单元数远多于 VisNet,这些都会影响定量结论的力度。

一句话总结

读下来我觉得这篇文章的说服力不在于 VisNet 赢了 SVM 分数——它其实没赢多少——而在于把"生物学合理性"从口号拆成了可操作的验收条款:不变性必须长在层级内的神经元里、稀疏到单细胞可读、对构型破坏敏感、用局部规则学会。HMAX 在这套验收下的溃败方式(相似度矩阵一片红、打乱面孔照放不误)尤其有画面感;当然作者自家模型的参数选择与"谁来定义合理"的标准之争,我也保留一分警惕(此段判断为我个人理解)。


审校与证据追溯 (Verification & Evidence)

图表审计结果

  • Fig1: 提取质量 good,对齐度 full,识别面板 []
  • Fig10: 提取质量 good,对齐度 full,识别面板 []
  • Fig11: 提取质量 good,对齐度 full,识别面板 []
  • Fig12: 提取质量 good,对齐度 full,识别面板 []
  • Fig13: 提取质量 good,对齐度 full,识别面板 []
  • Fig14: 提取质量 good,对齐度 full,识别面板 []
  • Fig2: 提取质量 good,对齐度 full,识别面板 []
  • Fig3: 提取质量 good,对齐度 full,识别面板 []
  • Fig4: 提取质量 good,对齐度 full,识别面板 []
  • Fig5: 提取质量 good,对齐度 full,识别面板 []
  • Fig6: 提取质量 good,对齐度 full,识别面板 []
  • Fig7: 提取质量 good,对齐度 full,识别面板 []
  • Fig8: 提取质量 good,对齐度 full,识别面板 []
  • Fig9: 提取质量 good,对齐度 full,识别面板 []

关键事实与局限性声明

  • 补充要点: 实验 2 的统计对照细节:未训练 VisNet 以模式联想解码得 18%(原文明确称为 statistical control),笔记已提及,但可补充这是针对"随机权重+强读出器即可分类"这一已知陷阱的对照(§2.8.2 末段),方法论意义重要
  • 补充要点: Booth & Rolls (1998) 的对照实验(未以此方式观看过的物体不产生视角不变表征)在原文 §4.5 被引为因果链关键一环,笔记讨论中引用该研究时未提及此对照
  • 补充要点: 图 9 原文图注存在明显笔误("The test images were 20° away from the test images",应为 training images),笔记图注解读已按正确含义转述但未标注原文笔误
  • 补充要点: HMAX VTU 路线的补充信息:原文指出 VTU 在近年 HMAX 工作中已不常用、末端 C 层直接送 SVM(p.523),可强化笔记关于"VTU 是补救措施"的论述