Discovery of input/output Casual relations in Software Systems
Discovery of input/output Casual relations in Software Systems
批准号:
2399261
负责人:
金额:
$0.0万
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
尽管数字技术显著地改变了日常生活,但安全和隐私问题也随之出现。他们任务的复杂性使得自动化系统越来越难以理解,导致信心的逐渐丧失,尤其是在涉及敏感数据的关键场景中。当前的可解释性方法仅从局部角度解决不透明性问题:适合于特定系统,突出统计相关性,或提供过于技术性的解释。我们的研究集中在因果发现领域,目的是回答“为什么一个系统会做出这样的决定?”这个关键问题,将分析提升到更高的抽象层次。我们提出了一种新的通用因果解释方法,用于大规模测试软件系统的行为逻辑,具有双重目的:根据直接影响它的输入类别,为自动化过程产生结果的原因提供人类可理解的解释;评估系统的决策逻辑是否违反了预定义的软件规范。我们利用信息论和在被测系统(SUT)的输入空间上定义的分区的点阵结构,将其视为黑盒:我们只需要访问输入/输出(I/O)接口。我们开发了一种算法,通过条件互信息(CMI)执行条件独立性测试来调查I/O因果关系。这个想法是寻找输入变量(输入部分)的最小子集,当它被改变时,会导致输出(或给定的输出部分)发生变化。使用消除策略来构建解释,排除具有null CMI的输入变量。本文研究的三大核心支柱是:1)基于信息理论的一般因果解释框架:不同应用领域的方法论评价。我们展示了我们的方法在测试软件和不同性质的属性方面的通用性。我们开始专注于基于机器学习的预测系统,涉及敏感输入类别,实施安全策略的程序和图像识别系统。在讨论了黑盒场景中的公平、信息泄露和错误分类之后,我们将介绍第四个白盒案例研究。随着越来越多的研究集中在测试社交网络平台上,我们发现采用我们的解释方法通过利用因果推理来模拟模拟的社交互动是很有趣的。2)增加统计保证,提高软件测试性能。我们的目标是通过为我们的测试方法提供统计保证来提高测试集和信息论测量的质量。我们调查了用于估计互信息的现有统计方法,以寻找最适合我们研究背景的方法,这些方法具有良好的统计置信度。这将允许:提供更严格的发现,证明检测到的有影响和无影响的部分是正确的,具有高可信度;从近似值和测试套件大小发布我们的发现。3)用有向信息(DI)概括因果分析。我们讨论了格兰杰因果关系和DI理论之间的联系,表明DI代数结构,简化到我们的场景,与CMI一致。因此,使用DI扩展我们的一般化方法成为我们研究的一个引人注目的方向,特别是在I/O关系方向不明确的更复杂的场景中。交互式程序(网页、编辑器)提供了一个有趣的应用程序上下文,其中信息可以向不同方向流动,假设存在多个用户同时与程序交互。我们期望为用户和开发人员提供广泛的贡献,使关键软件系统的知情使用,检测和减轻对公平和安全的威胁,并提高整体质量
英文摘要
Despite digital technologies have notably revolutionized daily life, security and privacy concerns have emerged alongside. The complexity of their tasks has made automated systems ever more difficult to understand, causing a progressive loss of confidence, especially in critical scenarios involving sensitive data. Current interpretability methods tackle the opacity problem only from partial perspectives: suited to specific systems, highlighting statistical correlations, or providing too technical explanations.Our research places in the Causal Discovery area and aims thus to answer the critical question "why did a system make that decision?" lifting the analysis to higher levels of abstraction. We propose a novel General Causal Explanation Method for testing the behavioural logic of software systems on large-scale, serving a dual purpose: provide human-understandable explanations of the reasons why an automated procedure comes to the outcome in terms of input categories that directly affect it; assess whether a system's decision logic violates predefined software specifications.We leverage on Information Theory and a Lattice structure of partitions defined on the input space of the systems under test (SUT), treated as black-box: we only require access to input/output (I/O) interfaces. We develop an algorithm that investigates I/O causal relationships by performing Conditional Independence Testing via Conditional Mutual Information (CMI). The idea is look for the smallest subset of input variables (input part) that when altered cause a change in the output (or given output part). An elimination strategy is used to build explanations, excluding input variables with null CMI.The three core pillars of our research are summarized below:1) Information-Theory grounded General Causal Explanation Framework: evaluation of the methodology in different application areas. We show the versatility of our method in testing software and properties of different nature. We start focusing on Machine Learning-based predictive systems involving sensitive input categories, Programs implementing security policies and Image recognition systems. After dealing with Fairness, Information leakage and misclassification in black-box scenarios, we'll introduce a fourth white-box case study. With the growing number of studies focused on testing Social Network platforms, we find interesting to adopt our explanation methodology to model simulated social interactions by leveraging on causal reasoning. 2) Enhance Software Testing Performance adding Statistical Guarantee.We aim to improve our Test Set and information-theoretic measurements' quality by providing our testing approach with statistical guarantee. We investigate existing statistical methods used for estimating Mutual Information looking for the most suitable in the context of our study, that comes with a good statistical confidence level. This would allow to: provide more rigorous findings, argue that the detected influential and non-influential parts are correct with a high-level confidence; release our findings from approximations and test suite size.3) Generalise Causal Analysis with Directed Information (DI).We discuss the link between Granger Causality and DI Theory showing that DI algebraic structure, simplified to our scenarios, aligns with CMI. Expand our generalised approach using DI becomes thus a compelling direction for our research, especially w.r.t. more complex scenarios where the direction of the I/O relationship is ambiguous. An interesting application context is given by interactive programs (webpages, editors) where the information can flow in different directions, given the existence of multiple users interacting with the program simultaneously. We expect to deliver a broad range of contributions both to users and developers, enabling informed usage of critical software systems, detecting and mitigating threats to fairness and security, and enhancing the overall q
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
近空间飞行器载MIMO SAR高分辨率、宽测绘带遥感成像机理与方法
-
批准号:41101317
-
项目类别:青年科学基金项目
-
资助金额:25.0万元
-
批准年份:2011
-
负责人:王文钦
-
依托单位: