Viewpoints on the History of Digital Synthesis

Viewpoints on the History of Digital Synthesis
复制标题

数字合成历史的观点

DOI:
--
复制
发表时间:
1991
期刊:
International Conference on Mathematics and Computing
影响因子:
--
通讯作者:
J. Smith
J. Smith
中科院分区:
--
文献类型:
--
作者:
J. Smith

文献摘要

被引文献

相似文献

由于缺乏分析支持,算法综合的重要性似乎注定会降低。正如许多算法合成尝试很久以前向我们展示的那样,通过探索数学表达式的参数来找到各种令人愉悦的音乐声音是很困难的。除了可能赋予某些意义的音乐背景之外,大多数声音根本就没意思。获得有趣声音的最直接方法是利用过去的乐器技术或自然声音。频谱建模和物理建模合成技术都可以对此类声音进行建模。在这两种情况下,模型都是由分析程序确定的,该分析程序计算最佳模型参数以近似特定的输入声音。音乐家操纵参数来创造音乐变化。获得对采样合成的更好控制将需要更通用的声音转换。为了实现这一目标,必须根据我们所听到的来理解转变。我们知道理解声音变换的最好方法是研究它对短时频谱的影响,其中频谱分析参数被调整为尽可能匹配听力特征。因此,采样合成似乎不可避免地会转向频谱建模。如果抽象方法消失,采样合成被吸收到谱建模中,那么就只剩下两类:物理建模和谱建模。这将所有合成技术归结为对声音源或接收器进行建模的技术。下表列出了每个案例的一些特征:“数字合成历史的观点”,J.O.史密斯,Proc。国际。计算机音乐会议(ICMC-91),蒙特利尔,第 1-10 页,1991 年 10 月。 5 对未来的预测 第 13 页 频谱建模 物理建模 完全通用 特殊情况具体分析 任何基底膜天际线 任何有一定成本的仪器 时域和频域 时域和空间域 大量时频包络 多个物理变量 内存要求 较大 更紧凑的描述 大操作数/样本 从小到大的复杂性 随机部分最初很容易 随机部分通常棘手 攻击 困难 攻击自然 发声困难 发声自然 表现力有限 表现力无限 非线性困难 非线性自然 延迟/混响困难 延迟/混响自然 可以校准到自然 可以校准到自然 可以校准到任何声音 可以校准到自己的声音 物理不太有用 物理非常有用 图片中的卡通 来自所有线索的工作模型 进化受限 进化无界 表示声音接收器 表示声源 因为频谱建模直接构建沿耳朵的基底膜,其范围本质上比物理建模的范围更广。然而,物理模型提供了更紧凑的算法来生成熟悉的声音类别,例如弦乐和木管乐器。此外,它们通常更有效地产生由攻击清晰度、长延迟、脉冲噪声或物理乐器中的非线性引起的频谱效果。停下来思考一下自音乐诞生以来表演音乐家如何与共鸣器互动也是很有趣的。当谐振器的脉冲响应持续时间大于频谱帧的脉冲响应持续时间(名义上是耳朵的“积分时间”)时(就像任何弦一样),直接在短时频谱中实现谐振器就会变得不方便。在短时频谱中,谐振器作为递归实现比作为超薄共振峰更容易实现。当然,正如 Orion Larson 所说:“在软件中一切皆有可能。”频谱建模在时域中存在未解决的问题:尚不知道如何在攻击或其他相敏瞬态附近最好地修改短时傅立叶分析。相位在瞬态期间很重要,但在稳态期间则不重要;适当的时变频谱模型应该仅在需要精确合成的地方保留相位。非平稳声音的音色感知的一般问题变得很重要。小波变换支持更通用的信号构建块,可以想象这些信号构建块可以帮助解决瞬态建模问题。迄今为止,大多数小波变换活动仅限于基本的恒定 Q 频谱分析,其中分析滤波器与对数频率网格对齐,并且带宽与中心频率或 Q 的比率恒定。频谱模型也还不复杂;正弦曲线和分段线性包络的过滤噪声是一个好的开始,但肯定还有其他好的原始私人通信,1976 年“数字合成历史的观点”,J.O.史密斯,Proc。国际。计算机音乐会议(ICMC-91),蒙特利尔,第 1-10 页,1991 年 10 月。
algorithm synthesis seems destined to diminish in importance due to the lack of analysis support. As many algorithmic synthesis attempts showed us long ago, it is difficult to find a wide variety of musically pleasing sounds by exploring the parameters of a mathematical expression. Apart from a musical context that might impart some meaning, most sounds are simply uninteresting. The most straightforward way to obtain interesting sounds is to draw on past instrument technology or natural sounds. Both spectral-modeling and physical-modeling synthesis techniques can model such sounds. In both cases the model is determined by an analysis procedure that computes optimal model parameters to approximate a particular input sound. The musician manipulates the parameters to create musical variations. Obtaining better control of sampling synthesis will require more general sound transformations. To proceed toward this goal, transformations must be understood in terms of what we hear. The best way we know to understand a sonic transformation is to study its effect on the short-time spectrum, where the spectrum-analysis parameters are tuned to match the characteristics of hearing as closely as possible. Thus, it appears inevitable that sampling synthesis will migrate toward spectral modeling. If abstract methods disappear and sampling synthesis is absorbed into spectral modeling, this leaves only two categories: physical-modeling and spectral-modeling. This boils all synthesis techniques down to those which model either the source or the receiver of the sound. Some characteristics of each case are listed in the following table: “Viewpoints on the History of Digital Synthesis,” J.O. Smith, Proc. Int. Computer Music Conf. (ICMC-91), Montreal, pp. 1-10, Oct. 1991. 5 PROJECTIONS FOR THE FUTURE Page 13 Spectral Modeling Physical Modeling Fully general Specialized case by case Any basilar membrane skyline Any instrument at some cost Time and frequency domains Time and space domains Numerous time-freq envelopes Several physical variables Memory requirements large More compact description Large operation-count/sample Small to large complexity Stochastic part initially easy Stochastic part usually tricky Attacks difficult Attacks natural Articulations difficult Articulations natural Expressivity limited Expressivity unlimited Nonlinearities difficult Nonlinearities natural Delay/reverb hard Delay/reverb natural Can calibrate to nature Can calibrate to nature Can calibrate to any sound May calibrate to own sound Physics not too helpful Physics very helpful Cartoons from pictures Working models from all clues Evolution restricted Evolution unbounded Represents sound receiver Represents sound source Since spectral modeling constructs directly the spectrum received along the basilar membrane of the ear, its scope is inherently broader than that of physical modeling. However, physical models provide more compact algorithms for generating familiar classes of sounds, such as strings and woodwinds. Also, they are generally more efficient at producing effects in the spectrum arising from attack articulations, long delays, pulsed noise, or nonlinearity in the physical instrument. It is also interesting to pause and consider how invariably performing musicians have interacted with resonators since the dawn of time in music. When a resonator has an impulse-response duration greater than that of a spectral frame (nominally the “integration time” of the ear), as happens with any string, then implementation of the resonator directly in the short-time spectrum becomes inconvenient. A resonator is a much easier to implement as a recursion than as a super-thin formant in a short-time spectrum. Of course, as Orion Larson says: “Anything is possible in software.” Spectral modeling has unsolved problems in the time domain: it is not yet known how to best modify a short-time Fourier analysis in the vicinity of an attack or other phase-sensitive transient. Phase is important during transients and not during steady-state intervals; a proper time-varying spectrum model should retain phase only where needed for accurate synthesis. The general question of timbre perception of non-stationary sounds becomes important. Wavelet transforms support more general signal building blocks that could conceivably help solve the transient modeling problem. Most activity with wavelet transforms to date has been confined to basic constant-Q spectrum analysis, where the analysis filters are aligned to a logarithmic frequency grid and have a constant ratio of bandwidth to center frequency or Q. Spectral models are also not yet sophisticated; sinusoids and filtered noise with piecewise-linear envelopes are a good start, but surely there are other good primiPrivate communication, 1976 “Viewpoints on the History of Digital Synthesis,” J.O. Smith, Proc. Int. Computer Music Conf. (ICMC-91), Montreal, pp. 1-10, Oct. 1991.