# \[Statistic\] Show Index number of Outliner in vector

**URL:** <https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172>\
**Category:** Data handling\
**Created:** [September 22, 2023, 1:27pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172 "2023-09-22T13:27:44Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Smart](https://avatars.discourse-cdn.com/v4/letter/s/8baadc/32.png) [@Smart](https://scilab.discourse.group/u/Smart)\
**Post date:** [September 22, 2023, 1:27pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/1 "2023-09-22T13:27:44Z")

</div>

Hello,

I have a matrix of measurement data, and one of the columns contains temperature measurements. Some of the data points contain outliers or anomalies. I’m currently using the Scilab Toolbox “samplestat” to filter out these outliers with the following function:

```auto
function [outlierfree, outlier] = ST_outlier(v, mod)

```

My goal is to obtain the indices of these “outliers.” I need these indices to remove the corresponding rows from the other columns in the matrix. For example, I want to align the good temperature values (outlierfree) with the time vector.

I’ve attempted to find all the outliers in the temperature vector using the `find()` function. However, there’s a possibility that the outlier value may occur multiple times in the temperature vector. Consequently, I end up with multiple indices for the same outlier value, and this could lead to the unintentional deletion of good data points.

Ideally, I’m looking for a function that can provide me with the indices of the outlier values.

I hope this clarifies my issue, and I’m hoping that someone can assist me with this. Best regards.

---

<div class="post-metadata">

**Author:** ![mottelet](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/mottelet/32/13_2.png) [@mottelet](https://scilab.discourse.group/u/mottelet)\
**Post date:** [September 22, 2023, 2:08pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/2 "2023-09-22T14:08:15Z")

</div>

Hello,

Please provide an example, with small matrices.

S.

---

<div class="post-metadata">

**Author:** ![Smart](https://avatars.discourse-cdn.com/v4/letter/s/8baadc/32.png) [@Smart](https://scilab.discourse.group/u/Smart)\
**Post date:** [September 22, 2023, 8:53pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/3 "2023-09-22T20:53:45Z")

</div>

Firstly, I want to express my gratitude for your prompt response. Your support serves as a great source of motivation for me.

Below, you will find a matrix:

If you plot column 8 against column 4, you will notice the presence of two outliers.

Now, let’s consider the use of the “Get Index in Plot” function:

```plaintext
% If I start with ST_outlier(Matrix(:,4)); , there will be no outliers found
[outlierfree, outlier] = ST_outlier(Matrix(10:end, 4)); 

--> outlier
outlier = 

   25.
   8.

% Obtain the index of outliers
for i = 1:size(outlier, '*')
    index(i) = find(Matrix(10:end, 4) == outlier(i));
end

% The issue arises here if the values of the outliers are found in the matrix, even if they are not outliers.

--> index = 

   11.
   15.

```

Now, what would be the most effective way to delete all rows corresponding to these indices? For example:

```plaintext
for i = 1:size(index)
     Matrix(index(i), :) = [];
end

```

Is it possible to delete all rows if I have used struct.row1, struct.row2, and so on?

I hope this example provides clarity on my predicament.

```auto
79.0	5.0	182.0	6.0	80.0	377.0	95.0	0.0
79.0	5.0	182.0	8.0	80.0	377.0	96.0	0.01
79.0	5.0	182.0	12.0	80.0	377.0	97.0	0.02
79.0	5.0	182.0	20.0	80.0	377.0	98.0	0.03
79.0	5.0	183.0	25.0	80.0	377.0	99.0	0.04
79.0	5.0	183.0	32.0	80.0	377.0	0.0	0.05
79.0	5.0	183.0	40.0	80.0	377.0	1.0	0.06
79.0	5.0	183.0	50.0	80.0	377.0	2.0	0.07
79.0	5.0	183.0	62.0	80.0	377.0	3.0	0.08
79.0	5.0	185.0	75.0	80.0	377.0	4.0	0.09
79.0	5.0	185.0	88.0	80.0	377.0	5.0	0.1
79.0	5.0	185.0	100.0	80.0	377.0	6.0	0.11
79.0	5.0	185.0	101.0	80.0	377.0	7.0	0.12
79.0	5.0	185.0	101.0	80.0	377.0	8.0	0.13
79.0	5.0	185.0	101.0	80.0	377.0	9.0	0.14
79.0	5.0	185.0	100.0	80.0	377.0	10.0	0.15
79.0	5.0	185.0	102.0	80.0	377.0	11.0	0.16
79.0	5.0	185.0	102.0	80.0	377.0	12.0	0.17
79.0	5.0	185.0	102.0	80.0	377.0	13.0	0.18
79.0	5.0	185.0	25.0	80.0	377.0	14.0	0.19
79.0	5.0	183.0	100.0	80.0	377.0	15.0	0.2
79.0	5.0	183.0	105.0	80.0	377.0	16.0	0.21
79.0	5.0	183.0	105.0	80.0	377.0	17.0	0.22
79.0	5.0	183.0	8.0	80.0	377.0	18.0	0.23
79.0	5.0	183.0	102.0	80.0	377.0	19.0	0.24
79.0	5.0	185.0	101.0	80.0	377.0	20.0	0.25
79.0	5.0	185.0	101.0	80.0	377.0	22.0	0.27
79.0	5.0	183.0	101.0	80.0	377.0	23.0	0.28
79.0	5.0	183.0	100.0	80.0	377.0	24.0	0.29
79.0	5.0	183.0	102.0	80.0	377.0	26.0	0.31
79.0	5.0	183.0	102.0	80.0	377.0	27.0	0.32
79.0	5.0	183.0	102.0	80.0	377.0	28.0	0.33
79.0	5.0	183.0	102.0	80.0	377.0	29.0	0.34
79.0	5.0	183.0	102.0	80.0	377.0	30.0	0.35
79.0	5.0	183.0	103.0	80.0	377.0	31.0	0.36
79.0	5.0	183.0	103.0	80.0	377.0	32.0	0.37
79.0	5.0	185.0	103.0	80.0	377.0	33.0	0.38
79.0	5.0	183.0	101.0	80.0	377.0	35.0	0.4
79.0	5.0	183.0	105.0	80.0	377.0	36.0	0.41
79.0	5.0	183.0	105.0	80.0	377.0	37.0	0.42
79.0	5.0	183.0	105.0	80.0	377.0	38.0	0.43

```

![image](https://global.discourse-cdn.com/free1/uploads/utc/original/1X/38421765e20bee76925d8707d304d9a44e86eebf.png)

Is it possible to extract the indices from the figure environment?

---

<div class="post-metadata">

**Author:** ![mottelet](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/mottelet/32/13_2.png) [@mottelet](https://scilab.discourse.group/u/mottelet)\
**Post date:** [September 23, 2023, 7:34pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/4 "2023-09-23T19:34:09Z")

</div>

Hello,

I understand the problem now. It will be faster to hack the ST\_outlier function in order to make it return the indexes of outliers…

S.

---

<div class="post-metadata">

**Author:** ![Smart](https://avatars.discourse-cdn.com/v4/letter/s/8baadc/32.png) [@Smart](https://scilab.discourse.group/u/Smart)\
**Post date:** [September 24, 2023, 11:16am UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/5 "2023-09-24T11:16:45Z")

</div>

Thank you, Stéphane. I will give it a try.

```auto
data1 = data(:, 4);

```

The results for ‘[of, o] = ST\_pearsonhartley(data1);’ and ‘[of, o] = ST\_outlier(data1);’ are identical.

I have managed to obtain the indices of the outliers directly. I made a minor modification to the code in ‘ST\_personhartley.sci’:

```auto
function [outlierfree, outlier, outlier_idx] = ST_pearsonhartley(v, p)

// Previous code

apifun_checklhs("ST_pearsonhartley", lhs, 1:3); // Output args

// Previous code

for i = 1:n
    q = abs((v(i) - X) / S);
    if (q > qcritval)
        outlier(j) = v(i);
        outlier_idx(j) = i; // Index of the outlier
        j = j + 1;

        continue;
    end
    outlierfree(k) = v(i);
    k = k + 1;
end

```

This modification works perfectly!

**However, there is still a question. How can I extend ‘qtable’ to handle a vector with more than 1000 values? But this will be solved after studying some statistic methods**

```auto
"Pearson-Hartley is only applicable for sample distributions with more than 3\n" + ..
"and fewer than 1000 values.")); 

```

```auto
/////////////////////////////////////////////////////////////////////////////////////////

```

I have also made changes to ‘ST\_outlier.sci’:  
I will certainly explore deeper if necessary, but for now, I am content with speeding up the process. Sometimes, you don’t need to be an expert in everything.

My modified version of ‘ST\_outlier.sci’ works perfectly, even though it is suggested using ‘ST\_personhartley.sci’. [GIT Hub scilab-samplestat](https://github.com/haniibrahim/scilab-samplestat/issues/1) Next weekend, I will attempt to understand the difference between these two functions, in order to unterstand the ‘ST\_personhartley.sci’ and the ‘qtable’ aspect.

Here is my current solution:

```auto
size(measurement) // Matrix size 1306 x 8
data = measurement(:, 4); // Only interested in finding outliers in column 4
[of, o, o_idx] = ST_outlier(data); // Find outliers and their indices in o_idx
measurement(o_idx, :) = []; // Remove all rows that match the indices in o_idx

```

---

<div class="post-metadata">

**Author:** ![mottelet](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/mottelet/32/13_2.png) [@mottelet](https://scilab.discourse.group/u/mottelet)\
**Post date:** [September 24, 2023, 9:19pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/6 "2023-09-24T21:19:23Z")

</div>

OK great. But I am quite doubtfull about the use of Pearson Hartley method on time series if they are non-stationary (cf. the example you gave above).

---

<div class="post-metadata">

**Author:** ![Smart](https://avatars.discourse-cdn.com/v4/letter/s/8baadc/32.png) [@Smart](https://scilab.discourse.group/u/Smart)\
**Post date:** [September 24, 2023, 9:59pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/7 "2023-09-24T21:59:13Z")

</div>

Hello once more, do you possess any expertise in this particular domain as well?  
What would you suggest for a dataset that is not stationary? I’ve actually found the ST\_outlier function to be quite effective. That’s one of the reasons I initiated the filtering process even when the dataset is in a stationary state.

---

<div class="post-metadata">

**Author:** ![mottelet](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/mottelet/32/13_2.png) [@mottelet](https://scilab.discourse.group/u/mottelet)\
**Post date:** [September 25, 2023, 7:18am UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/8 "2023-09-25T07:18:49Z")

</div>

I would use a local median filter to compute a smoothed version x\_f(k) of the signal x(k) and substract it from the original one to form the error e(k)=x(k)-x\_f(k), then compute the z-scores (e(k)-\mu)/\sigma (\mu is the mean error and \sigma its standard deviation), and finally apply a threshold of \pm 3 to find the indexes of the outliers (even \pm 2 should be ok since your signal is quite clean). Here is a possible implementation:

```auto
x=measurement(:,4);
t=1:size(x,1);
// filter with median filter of bandwidth 2*p+1
// value of p can be adjusted as needed.
p=2;
xf=x;
for i=p+1:length(x)-p
    xf(i)=median(x(i-p:i+p));
end

// compute the error
e=x-xf;

// compute z-scores and find outliers
s=stdev(e);
mu=mean(e);
z=(e-mu)/s;
k=find(abs(z)>3);
plot(t,x,t,xf,t,e);
plot(t(k),x(k),'o')
legend("$x$","$x_f$","error","outliers",-1)

```

 ![Screenshot 2023-09-25 at 11.18.05](https://global.discourse-cdn.com/free1/uploads/utc/original/1X/c05d64f7bd21ca53509d8d40ae5d0636061a1adc.png)

---

<div class="post-metadata">

**Author:** ![philipp](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/philipp/32/53_2.png) [@philipp](https://scilab.discourse.group/u/philipp)\
**Post date:** [September 29, 2023, 10:26am UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/9 "2023-09-29T10:26:38Z")

</div>

Mh,

if you have a smooth signal like this, you could also use:

d\_signal = diff(signal); // with signal beeing the data you are interested in

Background:  
At the position of the outliners, there is a big difference between conscutive data points.  
To be more precise: There is a big negative difference between conscutive data points.

I guess with a threshold value for the d\_signal it is possible to find the appropriate outliner positions.

Note:  
“d\_signal” has one less datapoint than “signal”, so you might need to adjust the index obtained by:  
index = find(d\_signal \< threshold); // with threshold beeing a negative value

Example:

x\_axis = Data(:,8);  
signal = Data(:,4);  
d\_signal = diff(signal);  
thresh = -20;  
index = find(d\_signal \< thresh);

// shift the index to the correct position  
index = index+1;

// plot the data  
plot(x\_axis, signal,‘-bl’);  
plot(x\_axis(index), signal(index),‘obl’);

![grafik](https://global.discourse-cdn.com/free1/uploads/utc/original/1X/9326dfd253e2b77150d113d401f770bc8b825407.png)

---

<div class="post-metadata">

**Author:** ![mottelet](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/mottelet/32/13_2.png) [@mottelet](https://scilab.discourse.group/u/mottelet)\
**Post date:** [October 2, 2023, 8:25am UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/10 "2023-10-02T08:25:14Z")

</div>

Hello,

Like you said, this method needs a smooth signal, but in addition would not be adequate if the signal is natively increasing or decreasing (and the threshold would not be easy to tune either). Substracting an average/mean allows to obtain a zero-mean signal on which you can apply classical statistics methods (e.g. z-scores) to detect outliers.

---

<div class="post-metadata">

**Author:** ![philipp](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/philipp/32/53_2.png) [@philipp](https://scilab.discourse.group/u/philipp)\
**Post date:** [October 4, 2023, 7:16pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/11 "2023-10-04T19:16:06Z")

</div>

Hello Stéphane,

you are absolutely right. The approach won’t work, if there are more than one outliners next to each other, since the difference between the outliners could be small.

---

<div class="post-metadata">

**Author:** ![philipp](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/philipp/32/53_2.png) [@philipp](https://scilab.discourse.group/u/philipp)\
**Post date:** [February 12, 2025, 4:10pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/12 "2025-02-12T16:10:15Z")

</div>

Although already a bit aged…I would like to continue this thread:

How would you find outliers that are close to the beginning and/or end of the dataset?

If I understand the following code correct values with index \< (p+1) and  
values with index \> (length(x)-p) are not taken into account.

hypothetical example:

- window width is 20, because there are 15 outliers next to each other.

The median calculation would start at index 21?  
Hence the first outlier would be corrected at index 21.  
Outliers with index \< window\_width would not be found…correct?

```auto
for i=p+1:length(x)-p
    xf(i)=median(x(i-p:i+p));
end

```

in other words:

```auto
e=x-xf;

```

is zero for the first and last value(s).

I guess: to remove these outliers one have to truncate the corrected data.

But what if one wants to keep the data size (nr. of samples)?

Best regards,  
Philipp

---

<div class="post-metadata">

**Author:** ![mottelet](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/mottelet/32/13_2.png) [@mottelet](https://scilab.discourse.group/u/mottelet)\
**Post date:** [February 13, 2025, 12:08pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/13 "2025-02-13T12:08:30Z")

</div>

Hello,

The best approach would be to use some extrapolation to add p leading and p trailing measurements to the original vector, then only consider the center `p+1:length(x)-p` points after filtering with the above approach. Maybe simple linear extrapolation would be the less risky, provided that first 2 points and last 2 measurements are not themselves outliers (and thus make the extrapolation irrelevant).

S.

---

<div class="post-metadata">

**Author:** ![philipp](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/philipp/32/53_2.png) [@philipp](https://scilab.discourse.group/u/philipp)\
**Post date:** [February 17, 2025, 8:55am UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/14 "2025-02-17T08:55:37Z")

</div>

Thank you.  
Best Regards,  
Philipp

---

<div class="post-metadata">

**Author:** ![mottelet](https://yyz2.discourse-cdn.com/free1/user_avatar/scilab.discourse.group/mottelet/32/13_2.png) [@mottelet](https://scilab.discourse.group/u/mottelet)\
**Post date:** [February 17, 2025, 4:33pm UTC](https://scilab.discourse.group/t/statistic-show-index-number-of-outliner-in-vector/172/15 "2025-02-17T16:33:37Z")

</div>

However, a median filter with a large window won’t be the best technique in your case (very long bursts of consecutive outliers). Maybe a model-based approach would work better (but I am not a specialist of these).

S.
