这是一篇综述,作者借用 David Marr 的"计算—表征—实现"三层框架,把人类腹侧颞叶皮层(ventral temporal cortex, VTC)的功能组织系统梳理了一遍,并提出了全文的核心假说:VTC 用一种"嵌套的空间层级"(nested spatial hierarchy)来安放类别信息——越抽象的信息铺在越大的皮层空间尺度上。它值得精读,是因为它把脑沟脑回这些"土办法"的解剖标志(尤其是中梭状回沟 mid-fusiform sulcus, MFS)与功能图谱严格对齐,为"为什么分类地图长成这个样子"给出了一个结构与功能绑定的解释框架。
研究背景
人类视觉分类快得惊人:大约十分之一秒内就能对场景做出归类,而这条通路要从初级视皮层(V1)出发,经过 V2、V3、人 V4 等一系列视网膜拓扑(retinotopic)区域,最后到达腹侧颞叶皮层。VTC 受损会造成各种失认症(agnosia),说明它确实是视觉识别的关键节点。此前的研究已经知道 VTC 里藏着大量信息——颜色、偏心距偏好(eccentricity bias)、视觉场地图、特定类别、专业知识、概念、语义、真实世界物体大小等,也发现了面孔、身体部位、场所、文字等类别选择性区域(如梭状回面孔区 FFA、海马旁回位置区 PPA、视觉词形区 VWFA)和跨体素的分布式表征。但当时的缺口在于:没有人能从计算的角度说清楚,VTC 究竟如何在解剖上组织这些信息、并借此实现高效分类。
还有一个更基础的困难:计算理论和表征研究都没有对"类别信息应该怎样铺在皮层表面"做出任何预测。换句话说,即便我们接受了 VTC 的功能,也依然不知道它的物理实现为什么要长成这样。这篇文章写作时恰好出现了一个新机会:VTC 的微观构筑(microarchitecture)、白质连接和宏观形态学(macroanatomy)研究刚取得进展,使得把结构数据与功能数据直接对齐、追问"结构如何服务于计算"第一次变得可行。作者于是以功能架构(functional architecture)为题,把这条 usually 缺失的链条补上。
研究思路
作者的总体策略是严格按 Marr 的三个层次推进:先问 VTC 的计算目标是什么,再看什么样的表征支持这些目标,最后追问这些计算与表征如何在皮层上物理实现。前两层用的是"理论要求—实证检验"的对照法:先把一个理想分类系统需要满足的条件(泛化、可分离性、多层级灵活读取)写清楚,再用 fMRI、颅内记录等数据逐条核对 VTC 是否满足。第三层则是这篇文章最有原创性的部分——他们注意到三类实现特征(聚类、拓扑组织、叠加),并把这些特征与具体的解剖标志联系起来。
在把三层各自梳理完后,作者再套用"整合系统观"(integrated systems perspective)把信息内容与物理实现连起来,提出核心假说:表征的空间尺度与其抽象层级挂钩,VTC 的空间层级为信息层级提供了神经基础设施。这个设计逻辑很清楚——如果结构与功能的对应在个体间可复现,就说明皮层硬件本身在为特定计算"定制"布局,那么顺着解剖去读功能就有了正当性。
方法
这是一篇综述,本身不做新实验,其"方法"体现为对多类证据的筛选与整合:人类功能磁共振(fMRI,含高分辨率成像与视网膜拓扑测量)、颅内脑电(intracranial recordings / electrocorticography,来自病人)、猕猴下颞叶(inferotemporal cortex, IT)的单细胞记录与 fMRI、死后组织的细胞构筑(cytoarchitectonics,如梭状回 FG1/FG2 分区)、受体构筑(receptor architectonics)、基于弥散的白质连接测量,以及人类连接组计划(HCP)约 196 名被试的组织对比度(被认为是髓鞘含量的代理指标)。
分析上的关键操作是"结构—功能对齐":在单个被试的膨胀皮层上,把各种大尺度功能地图(偏心距、领域特异性、animacy、真实世界物体大小)的内外侧转换位置,与 MFS、枕颞沟(OTS)、侧副沟(CoS)等解剖标志逐一对齐,检验转换边界是否可由脑沟预测。理解结果时需要注意,多数结论建立在这种"逐被试对齐"的证据之上,而非群体平均——作者自己也强调,平滑和群体平均会掩盖这种对应关系。
主要结果
- 计算目标可以明确写出并得到满足:VTC 表征既泛化又特异——对同一类别的不同格式(灰度、剪影)、不同范例保持更高反应,但对光照和视角变化敏感;位置、大小、镜像翻转则有一定容忍度。同时类别信息可分、且能按任务需要以上位、基本、下位多个抽象层级被提取(图 1)。
- VTC 而非早期视皮层持有"可读出"的类别信息:用简单线性分类器即可从 VTC 的分布式反应模式中准确解码类别,因为同类范例的分布式模式相似、异类相异;而 V1–V2 中同类之间的相似度甚至可能低于异类之间,说明"解缠"(untangling)发生在腹侧流后段(图 2b)。
- 分布式反应自带层级信息结构:VTC 的反应模式按上位(有生命/无生命)—基本(面孔/身体)—下位(人脸/动物脸)逐级聚类,且编码的是知觉相似性而非物理相似性,与人类行为判断一致(图 2c)。
- 三类实现特征成立:功能相似的神经元成簇(猕猴 fMRI 界定的簇内 29–97% 神经元偏好该类别,面孔簇最高达 56–97%);功能表征相对脑沟位置跨被试稳定(如 MFS 预测 pFus-faces/FFA-1 与 mFus-faces/FFA-2,OTS 预测身体部位区与 VWFA,CoS 预测 PPA);同一皮层区域上叠加多重表征(图 3)。
- MFS 是结构—功能对应的关键锚点:偏心距、领域特异性、animacy、真实世界物体大小四张地图的内外侧转换都对齐 MFS;MFS 同时分隔细胞构筑区 FG1(内侧、柱状)与 FG2(外侧、非柱状、细胞密度更高),受体构筑、髓鞘相关组织对比度和白质连接的转换也都落在 MFS 上(图 4)。
- 空间层级支持信息层级:整个 VTC 尺度(数厘米)承载上位的有生命/无生命区分,约 1 cm 尺度承载面孔、身体部位等生态类别,约 1 mm 的柱尺度承载眼睛、面孔朝向等具体特征——空间尺度越小,信息越具体,从而允许自上而下的门控机制按任务从不同尺度读出(图 5)。
图注解读
图 1 · 视觉分类系统的三大计算目标
原文图注:Figure 1 | Computational goals of a visual categorization system. a | The recognition system should generalize across a range of category exemplars — as well as across format and image transformations — while distinguishing between categories with similar features and configurations (for example, between faces of different species). b | To achieve efficient categorization, category information should be easy to read out. One way to achieve this efficiently is to have representations that are linearly separable. Assuming that an exemplar is represented by the distributed responses across a population of neurons, the computational constraint of separability entails that two exemplars of a category will evoke more similar distributed responses across the neural population than two exemplars of different categories (left graph). If this constraint is met, a simple linear classifier can be used to categorize stimuli (right graph). c | The recognition system should be able to extract several levels of information from a given input, as required by the task demands; in other words, it should enable flexible access to category information at several levels of abstraction.
这张图是全文的"需求说明书"。a 面板摆出第一对矛盾:系统既要对同类范例(以及格式、变换)泛化,又要在特征相近的类别之间保持特异性。b 面板解释什么叫可分离性:把每个范例看作神经元群体反应分布空间中的一个点,同类两点应当比异类两点靠得更近,这样一条虚线(线性分类器)就能把类别切开,实现快速且生物学上说得通的读出。c 面板用同一组输入示意多层级提取——同一对图片可以按"有生命/无生命"(上位)、"人脸/房子"(基本)、"男士/豪宅"(下位)乃至具体个人来解读。它支撑上文主要结果第 1 条:这三个目标是后文检验 VTC 表征与实现的基准。

图 2 · VTC 表征满足三大目标
原文图注:Figure 2 | Properties of the ventral temporal cortex representations. a | Generalization and specificity. Stronger functional MRI responses to faces are maintained across format (grey level and silhouettes) (left bar chart). Responses are higher for upright silhouettes than for upside-down silhouettes (right bar chart). **P < 0.001, significantly different from upright face silhouettes. Data from REF. 85. b | Separability of category information in the ventral temporal cortex (VTC) but not early visual cortex (V1–V2). Correlation matrices indicating the similarity between distributed responses to pairs of images from various categories (19 images per category) in the VTC and in V1–V2. In the top triangle, each cell shows the correlation between distributed responses to a pair of images. The bottom triangle shows the average correlation across images of a category. Hot colours indicate similar distributed response patterns and cold colours indicate dissimilar distributed response patterns. Data are from REF. 78 and show electrocorticography measurements in an example subject. c | Flexibility. Hierarchical clustering of distributed VTC responses measured with functional MRI reveals a separation between superordinate categories (inanimate versus animate), between basic-level categories (faces versus bodies) and between subordinate categories (human faces versus animal faces). This demonstrates that multiple levels of category information are represented in the VTC.
a面板用fMRI反应给出VTC表征泛化与特异性的神经证据。:对面孔区的 fMRI 反应在灰度图与剪影之间保持不变(左),但正立剪影显著高于倒置剪影(右,**P < 0.001)——泛化有边界,倒置破坏了构型信息。b 面板是相关矩阵,颜色越暖代表两图的分布式反应越相似:VTC 矩阵沿对角线呈亮块(同类相似、异类相异,类别信息可分),而 V1–V2 矩阵没有这种块状结构,说明可分离性是 VTC 的成就而非输入的自然属性。c 面板的层次聚类树先分出无生命/有生命,再分面孔/身体,再分人脸/动物脸,正是图 1c 所要求的灵活多层级读取在神经数据中的直接对应。这张图支撑主要结果第 2、3 条。

图 3 · 三大实现特征:聚类、拓扑、叠加
原文图注:Figure 3 | Three implementational features of the ventral temporal cortex: clustering, topological organization and superimposition. a | Neurons with similar category selectivity are clustered together. Each yellow dot indicates the location of a neuron that was recorded. In the enlarged version, red dots indicate individual face-selective neurons, and blue dots represent individual object-selective neurons in the macaque superior temporal sulcus (STS). b | Clustered functional regions responding to faces (red), places (green), words (brown), body parts (yellow) and objects (blue) have a consistent topology relative to macroanatomical landmarks in the human ventral temporal cortex (VTC). The mid-fusiform sulcus (MFS) predicts the location of the mid-fusiform face-selective region (mFus-faces (3); also known as FFA-2) and the posterior fusiform face-selective region (pFus-faces (2); also known as FFA-1). The inferior occipital gyrus (IOG) predicts the location of the IOG face-selective region (IOG-faces (1); also known as the occipital face area (OFA)). The occipitotemporal sulcus (OTS) predicts the location of both the occipitotemporal body part region (OTS-limbs (4); also known as the fusiform body area) and the visual word form area (VWFA (6)). The object-selective posterior fusiform/occipitotemporal sulcus (pFus/OTS (7)) partially overlaps with the VWFA and extends more posteriorly. The collateral sulcus (CoS) predicts the location of parahippocampal place area (PPA (5); also known as the CoS place-selective region (CoS-places)). As a result of these structure–function correspondences, there is a consistent topological organization among functional activations. For example, place-selective regions are medial to face-selective regions, whereas OTS-limbs separates pFus-faces from mFus-faces. Notably, within a given macroanatomical neighbourhood in the VTC, multiple representations are superimposed. For example, place-selective representations and retinotopic representations are superimposed along the CoS.
a 面板把猕猴 STS 中记录到的每个神经元画成一个点,放大后可见红色(面孔选择性)与蓝色(物体选择性)神经元并非随机混杂而是成簇——"聚类"在神经元尺度上成立。b 面板是 VTC 的功能分区图,每个颜色对应一个类别选择区,编号 1–7 标出各功能区的位置;读图要点是每个功能区都紧贴某条沟:IOG 预测 OFA,MFS 预测前后两个梭状回面孔区,OTS 预测身体部位区与 VWFA,CoS 预测 PPA。由此推出功能区之间的拓扑关系也是稳定的(位置区内侧于面孔区)。这张图支撑主要结果第 4 条:功能组织的"地址"可以由解剖形态预测。

图 4 · MFS:大尺度地图与微观构筑的共同边界
原文图注:Figure 4 | Linking anatomical features to large-scale functional maps in the ventral temporal cortex. a | The mid-fusiform sulcus (MFS) predicts transitions in many large-scale functional maps in the ventral temporal cortex (VTC). Lateral–medial functional transitions in the eccentricity bias map (based on data from REF. 17), the domain-specificity map (based on data from REF. 33), the animacy map (based on data from REF. 116) and the real-world object-size map (based on data from REF. 27 and T. Konkle, personal communication) are all aligned to the MFS (shown by the dashed black line). Each panel shows a representative inflated right hemisphere from an individual subject, with the exception of the domain-specificity map, which was generated from ten subjects. b | The MFS predicts transitions of anatomical features of the VTC. Lateral–medial anatomical transitions in cytoarchitecture, in white-matter connectivity, in the density of muscarinic acetylcholine receptor type 3 and in tissue contrast enhancement (which is thought to be related to myelin content) are each aligned to the MFS.
a 排四张功能地图(偏心距、领域特异性、animacy、真实世界物体大小),虚线黑线统一标出 MFS:四张图的内外侧转换都落在这条线上——中央凹/面孔/有生命/小物体在外侧,外周/场所/无生命/大物体在内侧。b 排对应的四张解剖图(细胞构筑 FG1/FG2、白质连接、M3 受体密度、被认为与髓鞘含量相关的组织对比度指标。)的转换同样对齐 MFS。这张图是全文最"硬"的证据:同一 条脑沟既是四张功能地图的边界,也是四种解剖属性的边界,支撑主要结果第 5 条,也是"结构约束功能"论点的核心依据。

图 5 · 嵌套的空间层级对应信息层级
原文图注:Figure 5 | The spatial structure of nested functional representations in the ventral temporal cortex supports the hierarchical information structure. a | Superimposition of functional representations in the ventral temporal cortex (VTC) from the animacy map (top) to clustered face-selective regions and body part-selective regions (middle) to clustering of neurons with shared response properties (bottom). b | Schematic hierarchy linking the spatial scale of functional representations implemented in the lateral VTC to the scale of information that each level represents. We propose that more-abstract information is represented at a larger spatial scale and more-concrete information at a finer spatial scale. We illustrate this idea with animate hierarchies as an example: superordinate information (animate) is represented at the scale of the entire VTC (several centimetres); information about ecological categories such as faces and body parts is represented at the centimetre scale; and exemplar information and complex-feature information is represented at the columnar level or an even smaller spatial scale. Additional hierarchies are likely to exist in the medial VTC and the VTC more generally.
a 面板从上到下做三个尺度的"套娃"展示:整块 VTC 的 animacy 地图、约 1 cm 的面孔区与身体部位区、以及共享反应性质的神经元柱。b 面板把这种套娃抽象成层级示意:上位信息(数厘米)—生态类别(厘米级)—复杂特征与范例信息(柱状尺度 100 μm–mm)。读图的关键是 b 中"空间尺度 ↔ 抽象程度"的对应箭头:越抽象越大尺度,越具体越小尺度。这张图是全文假说的落点,支撑主要结果第 6 条——VTC 的空间层级为多层级分类信息提供了可读出的物理基础。

讨论
作者的收束是三点结论:其一,尽管 VTC 表征空间的维度难以确定(一项基于电影刺激的数据挖掘估计在 35–50 维之间),其功能布局却出奇地有序——大尺度地图与细尺度簇彼此对齐,且都落在 MFS 两侧,形成一条跨表征维度共享的外侧—内侧功能梯度;这条梯度直接连到下层的连接与微观构筑。其二,空间层级生成了信息层级:把不同抽象层级安放在不同空间尺度上,可以提高类别加工的效率与灵活性,这一预测可以实证检验——干扰 VTC 不同空间尺度的活动应影响不同类型的分类决策,也可以用计算模型对照检验(带实现特征的模型 vs 不带的)。其三,叠加的表征通过不同程度的汇聚(convergence)与分岔(divergence)实现信息整合与分离:恢复限定语:“分岔可能支持独立信息的并行加工;FG1/FG2的差异微构筑可能反映针对不同计算的专门化。”,汇聚把相关信息放在一起、加速通信并降低布线成本,还允许多维信息的因子化组合以支持按任务灵活读取。
作者坦承的局限包括:结构—功能对应并非 1:1——大尺度地图大于细胞构筑区、细胞构筑区又大于单个功能簇,因此现有分区之下应该还有未发现的更细分区,每个细胞构筑区内也应包含多个功能簇;此外,一个体素(2–3 mm)内部多重表征如何排布(分层、跨柱还是到单个神经元)目前完全未知。对话对象上,这篇文章向上承接 Zeki & Shipp 与 Van Essen 等关于皮层汇聚/分岔连接逻辑的经典理论,横向与 Haxby 的分布式表征路线(聚类区 vs 分布式并不对立)和解剖学传统(Caspers 等的细胞构筑、Saygin 的连接组预测功能)呼应;作者还指出同样的"沟回预测功能"已在 V1、hV4、hMT+ 等处被观察到,提示这是皮层的普遍解法而非 VTC 特例。
一句话总结
在我看来,这篇文章真正的贡献不是发现某个新区域,而是用 MFS 这条此前解剖图谱里都没名字的脑沟做锚点,把散乱的功能区拉成一条"外侧—内侧"轴,再大胆地把"信息多抽象"与"皮层尺度多大"绑在一起,从而让 Marr 框架的第三层(实现)第一次有了可检验的形状。我自己读下来的判断是:尺度—抽象对应作为定性框架很有说服力,但作者自己也承认比例不是 1:1,且该假说对文字、场景等其余类别维度是否同样成立,文中留待未来验证——把它当作组织数据的坐标系,比当作定律更稳妥。
审校与证据追溯 (Verification & Evidence)
图表审计结果
- Fig1: 提取质量
good,对齐度full,识别面板[] - Fig2: 提取质量
good,对齐度full,识别面板[] - Fig3: 提取质量
good,对齐度full,识别面板[] - Fig4: 提取质量
good,对齐度full,识别面板[] - Fig5: 提取质量
good,对齐度full,识别面板[]
关键事实与局限性声明
- 审校纠偏: {'issue_id': 'ISSUE_002', 'current_level': '将候选机制和待检验预测写成已证实的空间—信息层级机制', 'supported_level': '作者提出的理论假说,与现有跨尺度组织证据一致但尚待因果和计算检验'} -> 评注:
- 审校纠偏: {'issue_id': 'ISSUE_003', 'current_level': '跨个体结构—功能相关证明硬件为计算定制', 'supported_level': '相关关系提示底层硬件和连接可能约束或优化功能组织'} -> 评注:
- 审校纠偏: {'issue_id': 'ISSUE_004', 'current_level': '空间对齐证明结构约束功能', 'supported_level': '跨研究空间对齐为结构约束假说提供相关性支持,因果方向未确定'} -> 评注:
- 审校纠偏: {'issue_id': 'ISSUE_007', 'current_level': '分岔生成专门化硬件并实现并行计算', 'supported_level': '分岔和微构筑差异可能支持这些计算优势'} -> 评注:
- 补充要点: : (第 6 页)
- 补充要点: : (第 6 页)
- 补充要点: : (第 5 页)
- 补充要点: : (第 3 页)