Untitled record
Abstract
Systems and methods include processors for receiving training data for a user activity; receiving bias criteria; determining a set of model parameters for a machine learning model including: (1) applying the machine learning model to the training data; (2) generating model prediction errors; (3) generating a data selection vector to identify non-outlier target variables based on the model prediction errors; (4) utilizing the data selection vector to generate a non-outlier data set; (5) determining updated model parameters based on the non-outlier data set; and (6) repeating steps (1)-(5) until a censoring performance termination criterion is satisfied; training classifier model parameters for an outlier classifier machine learning model; applying the outlier classifier machine learning model to activity-related data to determine non-outlier activity-related data; and applying the machine learning model to the non-outlier activity-related data to predict future activity-related attribut

Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
20 claims: 12 independent, 8 dependent
- 1عناصر الحماية 1. طريقة تشتمل على:استقبال- بواسطة معالج واحد على الأقل من جهاز حاسوبي واحد على الأقل مرتبط ببيئة إنتاج واحدة على الأقل- طلب نموذج جاهز للإنتاج يشتمل على مجموعة بيانات تدريب من سجلات البيانات؛ 5 حيث يشتمل كل سجل بيانات على متغير مستقل ومتغير مستهدف؛ تحديد- بواسطة المعالج الواحد على الأقل- مجموعة من معلمات النموذج لنموذج تعلم آلي واحد على الأقل يشتمل على: )1( تطبيق- بواسطة المعالج الواحد على الأقل- نموذج التعلم الآلي الواحد على الأقل الذي يتضمن مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من 10 قيم النموذج المتنبأ بها؛ )2( توليد- بواسطة المعالج الواحد على الأقل- مجموعة أخطاء من أخطاء عنصر البيانات عن طريق مقارنة مجموعة قيم النموذج المتنبأ بها بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )3( توليد- بواسطة المعالج الواحد على الأقل- متجه اختيار بيانات لتحديد المتغي ارت المستهدفة غير الشاذة استنادًا جزئيًا على الأقل إلى مجموعة الأخطاء الخاصة بأخطاء عنصر البيانات 15 ومعيار انحياز واحد على الأقل؛ )4( تطبيق- بواسطة المعالج الواحد على الأقل- متجه اختيار البيانات على مجموعة بيانات التدريب لتوليد مجموعة بيانات غير شاذة؛ )5( تحديد- بواسطة المعالج الواحد على الأقل- مجموعة من معلمات النموذج المحدثة لنموذج التعلم الآلي الواحد على الأقل استنادًا إلى مجموعة البيانات غير الشاذة؛ و 20 )6( تكارر- بواسطة المعالج الواحد على الأقل- تكارر واحد على الأقل للخطوات )1(-)5( حتى يتم استيفاء معيار إنهاء أداء رقابة واحد على الأقل للحصول على مجموعة معلمات النموذج لنموذج التعلم الآلي الواحد على الأقل في صورة معلمات النموذج المحدثة، حيث يقوم كل تك ارر بإعادة توليد مجموعة القيم المتنبأ بها، مجموعة الأخطاء، متجه اختيار البيانات، ومجموعة البيانات غير الشاذة باستخدام مجموعة معلمات النموذج المحدثة كمجموعة معلمات النموذج 25 الأولية؛ و 18476 -136- إرسال- بواسطة المعالج الواحد على الأقل- نموذج تعلم آلي جاهز للإنتاج لنموذج التعلم الآلي الواحد على الأقل الذي يستند جزئيًا على الأقل إلى تك ارر واحد على الأقل للاستخدام في بيئة إنتاج واحدة على الأقل.
- 25 2. الطريقة وفقًا لعنصر الحماية 1، حيث تشتمل أيضًا على:اختيار- بواسطة المعالج الواحد على الأقل- نموذج تعلم آلي واحد على الأقل للتحليلات الشاذة يستند جزئيًا على الأقل إلى طلب نموذج جاهز للإنتاج؛ تحديد- بواسطة المعالج الواحد على الأقل- مجموعة من معلمات نموذج التحليلات الشاذة لنموذج التعلم الآلي الواحد على الأقل للتحليلات الشاذة تشتمل على: 10 )7( تطبيق- بواسطة المعالج الواحد على الأقل- نموذج التعلم الآلي الواحد على الأقل للتحليلات الشاذة الذي يتضمن مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبئة بالتحليلات الشاذة؛ و )8( توليد- بواسطة المعالج الواحد على الأقل- مجموعة خطأ تحليلات لأخطاء عنصر بيانات التحليلات الشاذة عن طريق مقارنة مجموعة قيم التنبؤ المتنبأ بها بالتحليلات الشاذة بالقيم الفعلية 15 المناظرة لمجموعة بيانات التدريب؛ )9( تك ارر- بواسطة المعالج الواحد على الأقل- الخطوتين )7( و)8( كجزء من تك ارر واحد على الأقل حتى يتم استيفاء معيار إنهاء أداء الرقابة الواحد على الأقل لنموذج التعلم الآلي الواحد على الأقل؛ و توصيل- بواسطة المعالج الواحد على الأقل- نموذج التعلم الآلي للتحليلات الشاذة بالجهاز 20 الحاسوبي الواحد على الأقل لاستخدامه في بيئة إنتاج واحدة على الأقل للتنبؤ باحتمالية حدوث أحداث شاذة.
- 3الطريقة وفقًا لعنصر الحماية 1، حيث يشتمل المتغير المستقل لكل سجل بيانات على حالة شبكة كهربائية؛ 25 حيث تشتمل حالة الشبكة الكهربائية على:وقت من اليوم، 18476 -137- تاريخ، الطقس، موقع، كثافة سكانية، أو 5 أي توليفة منها؛ حيث يشتمل المتغير المستهدف على طلب طاقة شبكة؛ و حيث يشتمل نموذج تعلم آلي واحد على الأقل على نموذج تعلم آلي للتنبؤ بطلب الطاقة واحد على الأقل تم تدريبه للتنبؤ بطلب طاقة الشبكة استنادًا جزئيًا على الأقل إلى حالات الشبكة الكهربائية اللاحقة. 10
- 4الطريقة وفقًا لعنصر لحماية 3، حيث تشتمل أيضًا على:اختيار- بواسطة المعالج الواحد على الأقل- نموذج تعلم آلي لطلب شديد واحد على الأقل يستند جزئيًا على الأقل إلى طلب نموذج جاهز للإنتاج؛ تحديد- بواسطة المعالج الواحد على الأقل- مجموعة من معلمات نموذج الطلب الشديد لنموذج 15 التعلم الآلي للطلب الشديد الواحد على الأقل التي تشتمل على: )7( تطبيق- بواسطة معالج واحد على الأقل- نموذج تعلم آلي لطلب شديد واحد على الأقل يحتوي على مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ به للطلب الشديد؛ و )8( توليد- بواسطة المعالج الواحد على الأقل- مجموعة أخطاء طلب شديد لعنصر بيانات الطلب 20 الشديد من خلال مقارنة مجموعة قيم النموذج المتنبأ بها للطلب الشديد بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )9( تك ارر- بواسطة المعالج الواحد على الأقل- الخطوتين )7(- )8( كجزء من تك ارر واحد الأقل حتى يتم استيفاء معيار إنهاء أداء الرقابة الواحد على الأقل لنموذج التعلم الآلي الواحد الأقل. 18476 -138- اتصال- بواسطة المعالج الواحد على الأقل- نموذج التعلم الآلي للطلب الشديد بالجهاز الحاسوبي الواحد على الأقل لاستخدامه في بيئة الإنتاج الواحدة على الأقل للتنبؤ باحتمالية طلب الشبكة الشديد.
- 55 5. الطريقة وفقًا لعنصر الحماية 1، حيث يشتمل المتغير المستقل لكل سجل بيانات على خصائص المستخدم؛ حيث تشتمل خصائص المستخدم على:المتصفح، الموقع، 10 العمر، أو أي توليفة منها؛ حيث يشتمل المتغير المستهدف على: مصدر المحتوى، موقع المحتوى على صفحة الويب، 15 منطقة شاشة المحتوى، نوع المحتوى، التصنيف، أو أي توليفة منها؛ و حيث يشتمل نموذج التعلم الآلي الواحد على الأقل على نموذج تعلم آلي للتنبؤ بالمحتوى واحد على 20 الأقل تم تدريبه للتنبؤ بتوصية محتوى تستند جزئيًا على الأقل إلى خصائص المستخدم اللاحقة.
- 6طريقة تشتمل على:إرسال، بواسطة معالج واحد على الأقل لجهاز حاسوبي واحد على الأقل مرتبط ببيئة إنتاج واحدة على الأقل، طلب نموذج جاهز للإنتاج يشتمل على مجموعة بيانات تدريب من سجلات البيانات 25 إلى معالج توليد نموذج آلي واحد على الأقل؛ حيث يشتمل كل سجل بيانات على متغير مستقل ومتغير مستهدف؛ 18476 -139- استقبال، بواسطة المعالج الواحد على الأقل من معالج توليد نموذج آلي واحد على الأقل، نموذج تعلم آلي جاهز للإنتاج يستند جزئيًا على الأقل إلى تك ارر واحد على الأقل يتم إج ارؤه بواسطة معالج توليد النموذج الآلي الواحد على الأقل، حيث يشتمل التكارر الواحد على الأقل على: تحديد مجموعة من معلمات النموذج لنموذج التعلم الآلي الواحد على الأقل، حيث تشتمل على: 5 )1( تطبيق نموذج التعلم الآلي الواحد على الأقل الذي يتضمن مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها؛ )2( توليد مجموعة الأخطاء لأخطاء عنصر البيانات من خلال مقارنة مجموعة قيم النموذج المتنبأ بها بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )3( توليد متجه اختيار بيانات لتحديد المتغي ارت المستهدفة غير الشاذة استنادً ا جزئيًا على الأقل 10 إلى مجموعة الأخطاء الخاصة بأخطاء عنصر البيانات معيار انحياز واحد على الأقل؛ )4( تطبيق متجه اختيار البيانات على مجموعة بيانات التدريب لتوليد مجموعة بيانات غير شاذة؛ )5( تحديد مجموعة من معلمات النموذج المحدثة لنموذج التعلم الآلي الواحد على الأقل استنادًا إلى مجموعة البيانات غير شالاذة؛ و )6( تك ارر تك ارر واحد على الأقل للخطوات )1(-)5( حتى يتم استيفاء معيار إنهاء أداء رقابة 15 واحد على الأقل للحصول على مجموعة معلمات النموذج لنموذج التعلم الآلي الواحد على الأقل في صورة معلمات النموذج المحدثة، حيث يقوم كل تك ارر بإعادة توليد مجموعة القيم المتنبأ بها، مجموعة الأخطاء، متجه اختيار البيانات، ومجموعة البيانات غير الشاذة باستخدام مجموعة معلمات النموذج المحدثة في صورة مجموعة معلمات النموذج الأولية.
- 720 7. الطريقة وفقًا لعنصر الحماية 6، حيث تشتمل أيضًا على استقبال- بواسطة المعالج الواحد على الأقل- نموذج تعلم آلي لتحليلات شاذة للاستخدامه في بيئة الإنتاج الواحدة على الأقل للتنبؤ باحتمالية وقوع أحداث شاذة؛ حيث يشتمل التك ارر الواحد على الأقل أيضًا على:اختيار نموذج تعلم آلي لتحليلات شاذة واحد على الأقل استنادًا جزئيًا على الأقل إلى طلب نموذج 25 جاهز للإنتاج؛ 18476 -140- تحديد مجموعة من معلمات نموذج التحليلات الشاذة لنموذج التعلم الآلي للتحليلات الشاذة الواحد على الأقل الذي يشتمل على: )7( تطبيق نموذج التعلم الآلي للتحليلات الشاذة الواحد على الأقل الذي يتضمن مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها 5 للتحليلات الشاذة؛ و )8( توليد مجموعة أخطاء تحليلات شاذة لأخطاء عنصر بيانات تحليلات شاذة من خلال مقارنة مجموعة قيم النموذج المتنبأ بها للتحليلات الشاذة بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ و )9( تك ارر الخطوتين )7(- )8( كجزء من تك ارر واحد على الأقل حتى يتم استيفاء معيار إنهاء 10 أداء الرقابة الواحد على الأقل لنموذج التعلم الآلي الواحد على الأقل.
- 8الطريقة وفقًا لعنصر الحماية 6، حيث يشتمل المتغير المستقل لكل سجل بيانات على حالة شبكة كهربائية؛ حيث تشتمل حالة الشبكة الكهربائية على:15 وقت من اليوم، تاريخ، طقس، موقع، كثافة سكانية، أو 20 أي توليفة منها؛ حيث يشتمل المتغير المستهدف على طلب طاقة شبكة؛ و حيث يشتمل نموذج التعلم الآلي الواحد على الأقل على نموذج تعلم آلي لتنبؤ طلب الطاقة واحد على الأقل تم تدريبه للتنبؤ بطلب طاقة الشبكة استنادًا جزئيًا على الأقل إلى حالات الشبكة الكهربائية اللاحقة. 18476 -141-
- 9الطريقة وفقًا لعنصر الحماية 8، حيث تشتمل أيضًا على استقبال- بواسطة المعالج الواحد على الأقل- نموذج تعلم آلي لطلب شديد لاستخدامه في بيئة إنتاج واحدة على الأقل للتنبؤ باحتمالية طلب الشبكة الشديد؛ حيث يشتمل التك ارر الواحد على الأقل أيضًا على:5 اختيار نموذج تعلم آلي لطلب شديد واحد على الأقل استنادًا جزئيًا على الأقل إلى طلب نموذج جاهز للإنتاج؛ تحديد مجموعة من معلمات نموذج الطلب الشديد لنموذج التعلم الآلي للطلب الشديد الواحد على الأقل، حيث تشتمل على: )7( تطبيق نموذج التعلم الآلي للطلب الشديد الواحد على الأقل الذي يتضمن مجموعة من معلمات 10 النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها للطلب الشديد؛ و )8( توليد مجموعة أخطاء طلب شديد لأخطاء عنصر بيانات الطلب الشديد من خلال مقارنة مجموعة قيم النموذج المتنبأ بها للطلب الشديد بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ و )9( تك ارر الخطوتين )7(- )8( كجزء من تك ارر واحد على الأقل حتى يتم استيفاء معيار إنهاء 15 أداء الرقابة الواحد على الأقل لنموذج التعلم الآلي الواحد على الأقل.
- 10الطريقة وفقًا لعنصر الحماية 6، حيث يشتمل المتغير المستقل لكل سجل بيانات على خصائص المستخدم؛ حيث تشتمل خصائص المستخدم على:20 المتصفح، الموقع، العمر، أو أي توليفة منها؛ حيث يشتمل المتغير المستهدف على: 25 مصدر المحتوى، موقع المحتوى على صفحة الويب، 18476 -142- منطقة شاشة المحتوى، نوع المحتوى، التصنيف، أو أي توليفة منها؛ و 5 حيث يشتمل نموذج التعلم الآلي الواحد على الأقل على نموذج تعلم آلي للتنبؤ بالمحتوى واحد على الأقل تم تدريبه للتنبؤ بتوصية محتوى تستند جزئيًا على الأقل إلى خصائص المستخدم اللاحقة.
- 11نظام يشتمل على:معالج واحد على الأقل مهيأ لتنفيذ تعليمات الب ارمج التي تتسبب في قيام معالج واحد على الأقل 10 بتنفيذ الخطوات التالية: استقبال، من جهاز حاسوبي واحد على الأقل مرتبط ببيئة إنتاج واحدة على الأقل، طلب نموذج جاهز للإنتاج يشتمل على مجموعة بيانات التدريب لسجلات البيانات؛ حيث يشتمل كل سجل بيانات على متغير مستقل ومتغير مستهدف؛ تحديد مجموعة من معلمات النموذج لنموذج تعلم آلي واحد على الأقل تشتمل على: 15 )1( تطبيق نموذج التعلم الآلي الواحد على الأقل الذي يتضمن مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها؛ )2( توليد مجموعة أخطاء من أخطاء عنصر البيانات من خلال مقارنة مجموعة قيم النموذج المتنبأ بها بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )3( توليد متجه اختيار بيانات لتحديد المتغي ارت المستهدفة غير الشاذة استنادًا جزئيًا على الأقل 20 إلى مجموعة الأخطاء الخاصة بأخطاء عنصر البيانات ومعيار الانحياز الواحد على الأقل؛ )4( تطبيق متجه اختيار البيانات على مجموعة بيانات التدريب لتوليد مجموعة بيانات غير شاذة؛ )5( تحديد مجموعة من معلمات النموذج المحدثة لنموذج التعلم الآلي الواحد على الأقل استنادًا إلى مجموعة البيانات غير الشاذة؛ و )6( تك ارر تكارارً واحدًا على الأقل للخطوات )1(-)5( حتى يتم استيفاء معيار إنهاء أداء رقابة 25 واحد على الأقل للحصول على مجموعة معلمات النموذج لنموذج التعلم الآلي الواحد على الأقل في صورة معلمات النموذج المحدثة، حيث يعيد كل تك ارر توليد مجموعة القيم المتنبأ بها، مجموعة 18476 -143- الأخطاء، متجه اختيار البيانات، ومجموعة البيانات غير الشاذة باستخدام مجموعة معلمات النموذج المحدثة كمجموعة معلمات نموذج أولية؛ إرسال نموذج تعلم آلي جاهز للإنتاج لنموذج تعلم آلي واحد على الأقل يستند جزئيًا على الأقل إلى تك ارر واحد على الأقل للاستخدام في بيئة إنتاج واحدة على الأقل. 5
- 12النظام وفقًا لعنصر الحماية 11، حيث يتم تهيئة المعالج الواحد على الأقل بشكل إضافي لتنفيذ تعليمات الب ارمج التي تجعل المعالج الواحد على الأقل يقوم بتنفيذ الخطوات التالية:تحديد نموذج تعلم آلي لتحليلات شاذة واحد على الأقل استنادًا جزئيًا على الأقل إلى طلب نموذج جاهز للإنتاج؛ 10 تحديد مجموعة من معلمات نموذج التحليلات الشاذة لنموذج التعلم الآلي للتحليلات الشاذة الواحد على الأقل تشتمل على: )7( تطبيق نموذج التعلم الآلي للتحليلات الشاذة الواحد على الأقل الذي يتضمن مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم نموذج التحليلات الشاذة المتنبأ بها؛ و 15 )8( توليد مجموعة أخطاء للتحليلات الشاذة من أخطاء عنصر بيانات التحليلات الشاذة من خلال مقارنة قيم النموذج المتنبأ بها للتحليلات الشاذة بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )9( تك ارر الخطوتين )7(- )8( كجزء من تك ارر واحد على الأقل حتى يتم استيفاء معيار إنهاء أداء رقابة واحد على الأقل لنموذج التعلم الآلي الواحد على الأقل؛ و توصيل نموذج التعلم الآلي للتحليلات الشاذة بالجهاز الحاسوبي الواحد على الأقل لاستخدامه في 20 بيئة إنتاج واحدة على الأقل للتنبؤ باحتمالية حدوث أحداث شاذة.
- 13النظام وفقًا لعنصر الحماية 11، حيث يشتمل المتغير المستقل لكل سجل بيانات على حالة شبكة كهربائية؛ حيث تشتمل حالة الشبكة الكهربائية على:25 وقت من اليوم، تاريخ، 18476 -144- الطقس، موقع، كثافة سكانية، أو أي توليفة منها؛ 5 حيث يشتمل المتغير المستهدف على طلب طاقة شبكة؛ و حيث يشتمل نموذج التعلم الآلي الواحد على الأقل على نموذج تعلم آلي للتنبؤ بطلب الطاقة واحد على الأقل تم تدريبه للتنبؤ بطلب طاقة الشبكة استنادًا جزئيًا على الأقل إلى حالات الشبكة الكهربائية اللاحقة.
- 1410 14. النظام وفقًا لعنصر الحماية 13، حيث يتم تهيئة المعالج الواحد على الأقل أيضًا لتنفيذ تعليمات الب ارمج التي تتسبب في قيام المعالج الواحد على الأقل بتنفيذ الخطوات التالية:تحديد نموذج تعلم آلي للطلب الشديد واحد على الأقل استنادًا جزئيًا على الأقل إلى طلب نموذج جاهز للإنتاج؛ تحديد مجموعة من معلمات نموذج الطلب الشديد لنموذج التعلم الآلي للطلب الشديد الواحد على 15 الأقل تشتمل على: )7( تطبيق نموذج التعلم الآلي للطلب الشديد الواحد على الأقل الذي يحتوي على مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها للطلب الشديد؛ و )8( توليد مجموعة أخطاء طلب شديد من أخطاء عنصر بيانات الطلب الشديد من خلال مقارنة 20 مجموعة قيم النموذج المتنبأ بها للطلب الشديد بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )9( تك ارر الخطوتين )7(- )8( كجزء من تك ارر واحد على الأقل حتى يتم استيفاء معيار إنهاء أداء الرقابة الواحد على الأقل لنموذج التعلم الآلي الواحد على الأقل، توصيل نموذج التعلم الآلي للطلب الشديد بالجهاز الحاسوبي الواحد على الأقل للاستخدام في بيئة الإنتاج الواحدة على الأقل للتنبؤ باحتمالية طلب الشبكة الشديد. 25 18476 -145-
- 15النظام وفقًا لعنصر الحماية 11، حيث يشتمل المتغير المستقل لكل سجل بيانات على خصائص المستخدم؛ حيث تشتمل خصائص المستخدم على:المتصفح، 5 الموقع، العمر، أو أي توليفة منها؛ حيث يشتمل المتغير المستهدف على: مصدر المحتوى، 10 موقع المحتوى على صفحة الويب، منطقة شاشة المحتوى، نوع المحتوى، التصنيف، أو أي توليفة منها؛ و 15 حيث يشتمل نموذج التعلم الآلي الواحد على الأقل على نموذج تعلم آلي للتنبؤ بالمحتوى واحد على الأقل تم تدريبه للتنبؤ بتوصية محتوى تستند جزئيًا على الأقل إلى خصائص المستخدم اللاحقة.
- 16نظام يشتمل على:معالج واحد على الأقل لجهاز حاسوبي واحد على الأقل مرتبط ببيئة إنتاج واحدة على الأقل، حيث 20 يتم تهيئة المعالج الواحد على الأقل لتنفيذ تعليمات الب ارمج التي تتسبب في قيام المعالج الواحد على الأقل بتنفيذ الخطوات التالية: توصيل طلب نموذج جاهز للإنتاج يشتمل على مجموعة بيانات تدريب من سجلات البيانات بمعالج توليد نموذج آلي واحد على الأقل؛ حيث يشتمل كل سجل بيانات على متغير مستقل ومتغير مستهدف؛ 18476 -146- استقبال، من معالج توليد نموذج آلي واحد على الأقل، نموذج تعلم آلي جاهز للإنتاج متوقف جزئيًا على الأقل على تك ارر واحد على الأقل يتم إج ارؤه بواسطة معالج توليد النموذج الآلي الواحد على الأقل، حيث يشتمل التكارر الواحد على الأقل على: تحديد مجموعة من معلمات النموذج لنموذج التعلم الآلي الواحد على الأقل تشتمل على: 5 )1( تطبيق نموذج التعلم الآلي الواحد على الأقل الذي يتضمن مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها؛ )2( توليد مجموعة الأخطاء لأخطاء عنصر البيانات من خلال مقارنة مجموعة قيم النموذج المتنبأ بها بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )3( توليد متجه اختيار بيانات لتحديد المتغي ارت المستهدفة غير الشاذة استنادًا جزئيًا على الأقل 10 إلى مجموعة الأخطاء الخاصة بأخطاء عنصر البيانات ومعيار انحياز واحد على الأقل؛ )4( تطبيق متجه اختيار البيانات على مجموعة بيانات التدريب لتوليد مجموعة بيانات غير شاذة؛ )5( تحديد مجموعة من معلمات نموذج محدثة لنموذج التعلم الآلي الواحد على الأقل استنادًا إلى مجموعة البيانات غير الشاذة؛ و )6( تك ارر تك ارر واحد على الأقل للخطوات )1(-)5( حتى يتم استيفاء معيار إنهاء أداء رقابة 15 واحد على الأقل وذلك للحصول على مجموعة معلمات النموذج لنموذج التعلم الآلي الواحد على الأقل في صورة معلمات النموذج المحدثة، حيث يعيد كل تك ارر توليد مجموعة القيم المتنبأ بها، مجموعة الأخطاء، متجه اختيار البيانات، ومجموعة البيانات غير الشاذة باستخدام مجموعة معلمات النموذج المحدثة في صورة مجموعة معلمات نموذج أولية.
- 1720 17. النظام وفقًا لعنصر الحماية 16، حيث يتم تهيئة المعالج الواحد على الأقل لتنفيذ تعليمات الب ارمج التي تجعل المعالج الواحد على الأقل يقوم بتنفيذ الخطوات لاستقبال نموذج تعلم آلي للتحليلات الشاذة للاستخدام في بيئة إنتاج واحدة على الأقل للتنبؤ باحتمالية وقوع أحداث شاذة؛ حيث يشتمل التك ارر الواحد على الأقل أيضًا على:اختيار نموذج تعلم آلي للتحليلات الشاذة واحد على الأقل استنادًا جزئيًا على الأقل إلى طلب 25 نموذج جاهز للإنتاج؛ 18476 -147- تحديد مجموعة من معلمات نموذج التحليلات الشاذة لنموذج التعلم الآلي للتحليلات الشاذة الواحد على الأقل تشتمل على: )7( تطبيق نموذج التعلم الآلي للتحليلات الشاذة الواحد على الأقل الذي يحتوي على مجموعة من معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها 5 للتحليلات الشاذة؛ و )8( توليد مجموعة أخطاء تحليلات شاذة من أخطاء عنصر بيانات التحليلات الشاذة من خلال مقارنة مجموعة قيم النموذج المتنبأ بها للتحليلات الشاذة بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )9( تك ارر الخطوتين )7(- )8( كجزء من تك ارر واحد على الأقل حتى يتم استيفاء معيار إنهاء 10 أداء الرقابة الواحد على الأقل لنموذج التعلم الآلي الواحد على الأقل.
- 18النظام وفقًا لعنصر الحماية 16، حيث يشتمل المتغير المستقل لكل سجل بيانات على حالة شبكة كهربائية؛ حيث تشتمل حالة الشبكة الكهربائية على:15 وقت من اليوم، تاريخ، الطقس، موقع، كثافة سكانية، أو 20 أي توليفة منها؛ حيث يشتمل المتغير المستهدف على طلب طاقة شبكة؛ و حيث يشتمل نموذج التعلم الآلي الواحد على الأقل على نموذج تعلم آلي للتنبؤ بطلب الطاقة واحد على الأقل تم تدريبه للتنبؤ بطلب طاقة الشبكة استنادًا جزئيًا على الأقل إلى حالات الشبكة الكهربائية اللاحقة. 18476 -148-
- 19النظام وفقًا لعنصر الحماية 18، حيث يتم تهيئة المعالج الواحد على الأقل لتنفيذ تعليمات الب ارمج التي تجعل المعالج الواحد على الأقل يقوم بتنفيذ خطوات استقبال نموذج تعلم آلي لطلب شديد لاستخدامه في بيئة إنتاج واحدة على الأقل للتنبؤ باحتمال طلب شبكة شديد؛ حيث يشتمل التك ارر الواحد على الأقل أيضًا على:5 اختيار نموذج تعلم آلي لطلب شديد واحد على الأقل استنادًا جزئيًا على الأقل إلى طلب نموذج جاهز للإنتاج؛ تحديد مجموعة من معلمات نموذج الطلب الشديد لنموذج التعلم الآلي للطلب الشديد الواحد على الأقل تشتمل على: )7( تطبيق نموذج التعلم الآلي للطلب الشديد الواحد على الأقل الذي يحتوي على مجموعة من 10 معلمات النموذج الأولية على مجموعة بيانات التدريب لتحديد مجموعة من قيم النموذج المتنبأ بها للطلب الشديد؛ و )8( توليد مجموعة أخطاء طلب شديد من أخطاء عنصر بيانات الطلب الشديد من خلال مقارنة مجموعة قيم النموذج المتنبأ بها للطلب الشديد بالقيم الفعلية المناظرة لمجموعة بيانات التدريب؛ )9( تك ارر الخطوتين )7(- )8( كجزء من تك ارر واحد على الأقل حتى يتم استيفاء معيار إنهاء 15 أداء الرقابة الواحد على الأقل لنموذج التعلم الآلي الواحد على الأقل.
- 20النظام وفقًا لعنصر الحماية 16، حيث يشتمل المتغير المستقل لكل سجل بيانات على خصائص المستخدم؛ حيث تشتمل خصائص المستخدم على:20 المتصفح، الموقع، العمر، أو أي توليفة منها؛ حيث يشتمل المتغير المستهدف على: 25 مصدر المحتوى، موقع المحتوى على صفحة الويب، 18476 -149- منطقة شاشة المحتوى، نوع المحتوى، التصنيف، أو أي توليفة منها؛ و 5 حيث يشتمل نموذج التعلم الآلي الواحد على الأقل على نموذج تعلم آلي للتنبؤ بالمحتوى واحد على الأقل تم تدريبه للتنبؤ بتوصية محتوى تستند جزئيًا على الأقل إلى خصائص المستخدم اللاحقة. 18476 -150-
Independent claims20
1,653 paragraphs in 4 sections, as filed
Full Description
Background of the sister
Part of this patent disclosure contains copyrighted material. The copyright owner has no objection to the facsimile reproduction by any person of the patent document or patent disclosure as it appears in the files or records of the Patent and Trademark Office.
<p dir="rtl">5 Commercial, but otherwise reserves all copyrights whatsoever. The following notice applies to the software and data as shown below and in the graphics that are part of this document: Copyright, Hartford Steam Boiler Inspection and Insurance Company, All Rights Reserved.</p>
The present disclosure generally relates to improved computer-based systems, computer components and computer elements 10 configured to reduce bias in machine learning models.
A machine learning model may include one or more computers or processors to form predictions or determinations based on patterns and inferences learned from sample/training data. Bias in the selection of sample/training data can propagate into the predictions and determinations of a machine learning model.
General description of the invention
<p dir="rtl">15 The current disclosure models include low-bias dynamic anomaly machine learning models. The methods comprise receiving—by at least one processor—a training dataset of target variables representing an attribute associated with at least one activity of at least one user activity;</p>
18476
-3-
receiving - by at least one processor - at least one bias criterion used to identify one or more outliers; determining - by at least one processor - a set of model parameters for a machine learning model, wherein: (1) applying - by at least one processor - the machine learning model containing the set of initial model parameters to the training data set to identify a set of predicted model values; (2) generating - by at least one processor - a set of errors
(3) generating, by at least one processor, a data selection vector to identify non-anomous target variables based at least in part on the set of data element errors and at least one bias criterion; (4) using, by
<p dir="rtl">10 (5) selecting, by at least one processor, a data selection vector on the training dataset to generate a non-anomalous dataset; and (6) repeating, by at least one processor, steps (1)-(5) in a specified iteration until at least one control performance termination criterion is met to obtain the set of model parameters for the machine learning model as model parameters.</p>
<p dir="rtl">15 updated, where each iteration regenerates the set of predicted values, the set of errors, and the non-anomalous data set using the updated set of model parameters as the initial set of model parameters; training—by at least one processor—a set of classifier model parameters for a machine learning model of the anomalous classifier to obtain a trained machine learning model of the anomalous classifier that is initialized to identify at least one anomalous data item, based at least in part on the training data set</p>
<p dir="rtl">20 and data selection vector; applying, by at least one processor, the trained machine learning model of the anomaly classifier to a dataset of activity-related data for at least one user activity, to identify: 1) a set of anomalous activity-related data in the activity-related data set, and 2) a set of non-anomalous activity-related data in the activity-related data set; and applying, by at least one processor, the machine learning model to</p>
<p dir="rtl">25 A set of non-anomalous activity-related data elements that predicts the attribute of future activity relevant to at least one user's activity.</p>
18476
-4-
Embodiments of the present disclosure include systems for dynamically biased machine learning models. Systems having at least one processor connected to a non-transitory computer-readable storage medium with software instructions stored therein, wherein the software instructions, when executed, cause at least one processor to perform the following steps: receive a set of
<p dir="rtl">5 training data of target variables representing a characteristic associated with at least one activity of at least one user activity; receiving at least one user bias criterion to identify one or more outliers; determining a set of model parameters for a machine learning model including: (1) applying the machine learning model that includes a set of initial model parameters to the training data set to determine a set of predicted model values; (2) generating a set of errors from the element errors</p>
<p dir="rtl">10 (3) Generate a data selection vector to identify non-anomous target variables based on at least a partial set of data element errors and at least one bias criterion; (4) Use the data selection vector on the training data set to generate a non-anomous data set; (5) Identify a set of updated model parameters for the machine learning model based on the set of</p>
<p dir="rtl">15 Non-abnormal data; and (6) repeating steps (1)-(5) with a specified iteration until at least one control termination criterion is met to obtain the set of model parameters for the machine learning model in the updated model variables, where each iteration regenerates the set of predicted values, the error set, the data selection vector, and the non-abnormal data set using the updated set of model parameters as the initial model parameters; training a set of classifier model parameters for the machine learning model</p>
<p dir="rtl">20 For the anomalous classifier to have a trained machine learning model for the anomalous classifier initialized to identify an element</p>
at least one anomalous data based at least partially on the training dataset and the data selection vector; applying the trained machine learning model of the anomaly classifier to a dataset of activity-related data for at least one user activity to identify: 1) a set of anomalous activity-related data in the activity-related dataset, and 2) a set of non-anomalous data
<p dir="rtl">25 activity-related dataset activity-related dataset; and applying a machine learning model to</p>
18476
-5-
A set of non-anomalous activity-related data elements to predict the attribute of future activity relevant to at least one user's activity.
The systems and methods of the present disclosure models include: applying - by at least one processor - a data selection vector to a training data set to identify an anomalous training data set; training -
<p dir="rtl">5 By at least one processor—using the outlier training dataset, at least one outlier-specific model parameter for the at least one outlier-specific machine learning model to predict the outlier data values; and using—by at least one processor—the outlier-specific machine learning model to predict the activity-related outlier data values for the activity-related outlier dataset.</p>
Current detection models and methods include: Training - by at least one processor -
<p dir="rtl">10 Using the training dataset, general model parameters of a general machine learning model to predict data values; using, by at least one processor, the general machine learning model to predict activity-related anomalous data values for the activity-related anomalous dataset; and using, by at least one processor, the general machine learning model to predict activity-related data values.</p>
Systems and methods of embodiments of the present disclosure include: implementing - by at least one processor - a vector
<p dir="rtl">15 Selecting data on a training dataset to identify an outlier training dataset; training-by-at least one processor using the outlier training dataset, outlier-specific model parameters of an outlier-specific machine learning model to predict outlier data values; training-by-at least one processor using the training dataset, general model parameters of a general machine learning model to predict data values; using-by-at least one processor model</p>
<p dir="rtl">20 An outlier machine learning to predict activity-related anomaly values for an activity-related anomaly data set; and using, by at least one processor, an outlier machine learning model to predict activity-related anomaly values.</p>
The systems and methods of the present disclosure include: training - by at least one processor - using a training data set, general model parameters of a general machine learning model to predict data values;
18476
-6-
Using, by at least one processor, a general machine learning model to predict activity-related data values for the activity-related data set; using, by at least one processor, an anomaly classifier machine learning model to identify activity-related anomaly values for the activity-related data values; and removing, by at least one processor, the activity-related anomaly values.
<p dir="rtl">5 Systems and methods of the present disclosure models according to the present disclosure, wherein the training data set includes at least one activity-related attribute of concrete compressive strength as a function of concrete composition and exposure to concrete curing.</p>
Systems and methods of the present disclosure models according to the present disclosure, wherein the training dataset includes at least one activity-related attribute of energy use data as a function of household environmental conditions 10 and lighting conditions.
Systems and methods of embodiments of the present disclosure further include: receiving, by at least one processor, an application programming interface (API) request to generate a prediction of at least one data item; instantiating, by at least one processor, at least one cloud computing resource to schedule the execution of the machine learning model; using, by at least one processor in accordance with the execution schedule, the machine learning model
<p dir="rtl">15 Automated prediction of at least one activity-related data element value for at least one data element; and return, by at least one processor, the at least one activity-related data element value to a computing device associated with the API request.</p>
The systems and methods of the present disclosure include wherein the training dataset includes at least one activity-related feature of 3D patient images of a medical dataset; and wherein the machine learning model 20 is configured to predict activity-related data values that include two or more physical display parameters based on the medical dataset.
Systems and methods of the present disclosure include wherein the training data set includes at least one activity-related attribute of the simulated control results and the electronic machine command; and wherein the training data set includes at least one activity-related attribute of the simulated control results and the electronic machine command;
18476
-7-
Machine learning model to predict activity-related data values involving electronic control commands.
Systems and methods of embodiments of the present disclosure include: partitioning - by at least one processor - a set of activity-related data into a plurality of activity-related data subsets;
<p dir="rtl">5 Determining, by at least one processor, a consistent set model for each subset of the activity-related data; wherein the machine learning model includes a consistent set of models; wherein each consistent set model includes a random combination of models from the consistent set of models; using, by at least one processor, each consistent set model separately to predict values of the activity-related data for the consistent set; determining, by at least one processor</p>
<p dir="rtl">10 At a minimum, an error for each consistent set model based on the activity-related data values of the consistent set and the known values; and selecting - by at least one processor - the best performing consistent set model based on the lowest error.</p>
Brief explanation of the drawings
The various forms of the present disclosure can be further explained by reference to the attached drawings, wherein
<p dir="rtl">15 Similar structures are indicated by similar numbers in all of the multiple views. The drawings shown are not necessarily intended to scale, but are instead generally intended to illustrate the principles of the present disclosure. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but should be construed only as a representative basis for the illustration of which one skilled in the art may use one or more of the illustrative embodiments differently.</p>
<p dir="rtl">20 Figures 1-14B illustrate one or more schematic process flow charts, certain computer-based architectures, and/or screen shots of various specialized graphical user interfaces that illustrate at least some illustrative aspects of some embodiments of the present disclosure.</p>
Detailed description:
18476
-8-
Various detailed embodiments of the present disclosure are disclosed herein, in conjunction with the accompanying figures; however, it should be understood that the embodiments disclosed are merely illustrative. In addition, each of the examples provided in connection with the various embodiments of the present disclosure are intended to be illustrative and non-limiting.
<p dir="rtl">5 Throughout the entire description of the invention, the following terms shall have the meanings expressly assigned herein, unless the context clearly indicates otherwise. The phrases “in one embodiment” and “in some embodiments” as used herein do not necessarily refer to the same embodiment(s), although they may. Furthermore, the phrases “in another embodiment” and “in some other embodiments,” as used herein, do not necessarily refer to a different embodiment, although they may. Therefore,</p>
<p dir="rtl">10 As described below, many embodiments can be easily combined, without departing from the scope or intent of the present disclosure.</p>
In addition, the term "based on" is not exclusive and allows for reliance on additional factors not described, unless the context clearly indicates otherwise. In addition, throughout the description of the invention, singular words include plural references. The meaning of "in" includes "in"
<p dir="rtl">15 And "on".</p>
It is understood that at least one aspect/function of the various models described herein can be performed in real time and/or dynamically. As used herein, the term “real time” relates to an event/action that can occur instantaneously or nearly instantaneously at the time another event/action occurs. For example, “real-time processing,” “real-time computation,” and “real-time execution” all relate to
<p dir="rtl">20 “Actual” is defined as performing a computation in real time as it occurs for the relevant physical process (e.g., a user interacting with an application on a mobile device), so that the results of the computation can be used to drive the physical process.</p>
As used herein, the term “dynamically” and the term “automatically” and their derivatives and/or logical and/or semantically related terms mean that some event and/or action can be triggered and/or occur without any
18476
-9-
Human Intervention. In some embodiments, events and/or actions according to the present disclosure may be in real time and/or contingent on a predetermined periodicity of at least: nanosecond, several nanoseconds, millisecond, several milliseconds, second, several seconds, minute, several minutes, hourly, several hours, daily, several days, weekly, monthly, etc.
<p dir="rtl">5 In some embodiments, the illustrative inventive computer systems are specially programmed and configured with relevant hardware to operate in a distributed network environment, communicating with each other via one or more suitable data communications networks (e.g., the Internet, satellite, etc.) and using one or more of the most suitable data communications protocols/modes, including but not limited to, AppleTalk(TM), AX.25, X.25, IPX/SPX,</p>
<p dir="rtl">10 TCP/IP (e.g., HTTP), Near Field Communication (RFID, NFC), Narrowband Internet of Things (WiMax, WiFi, GSM, GPRS, NBIOT, CDMA, 3G, 4G, 5G), Satellite, ZigBee, and other appropriate communication modes. In some embodiments, NFC may represent a short-range wireless communication technology in which NFC-enabled devices are “snapped,” “squeezed,” “pressed,” or moved at close range to communicate.</p>
<p dir="rtl">15 The material disclosed herein may be implemented in software, firmware, or a combination thereof, or as instructions stored on a machine-readable medium, which can be read and executed by one or more processors. The machine-readable medium may include any medium and/or mechanism for storing or transmitting information in a form that can be read by a machine (e.g., a computer). For example, the machine-readable medium may include read-only memory (ROM); random access memory (RAM);</p>
<p dir="rtl">20 Magnetic disk storage media; optical storage media; flash storage media; electrical, optical, acoustic or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.) and others.</p>
As used herein, the terms "computer engine" and "engine" identify at least one software component and/or a combination of at least one software component and at least one hardware component.
18476
-10-
Designed/programmed/configured to manage/control software components and/or other hardware components (e.g., libraries, SDKs, objects, etc.).
Examples of hardware elements can include processors, microprocessors, circuits, circuit elements (such as transistors, resistors, capacitors, inductors, etc.), circuits
<p dir="rtl">5 Integrated circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate arrays (FPGA), logic gates, registers, semiconductor device, chips, microchips, chip stacks, etc. In some embodiments, one or more processors may be implemented as Complex Instruction Set Computer (CISC) or Reduced Instruction Set Computer (RISC) processors; 10 x86-compatible processors; multi-core microprocessor; or any other microprocessor or central processing unit (CPU). In various implementations, the one or more processors may be dual-core processor(s), dual-core mobile processor(s), etc.</p>
Examples of programs can include: program components, programs, applications, computer programs, application programs, system programs, machine programs, operating system programs, intermediate programs, 15 fixed programs, program modules, routines, subroutines, functions, methods,
Procedures, program interfaces, application programming interfaces (APIs), instruction sets, computer code, computer code, code fragments, computer code parts, words, values, symbols, or any combination thereof. Whether a model is implemented using hardware elements and/or software elements can vary depending on a number of factors, such as arithmetic rate, power levels, temperature tolerance, processing cycle budget, input data rates, output data rates, memory resources
Required data bus speeds and other design or performance constraints.
One or more aspects of at least one embodiment may be implemented by representational instructions stored on a machine-readable medium that represent various logic within the processor, which when read by a machine causes the machine to synthesize the logic of the techniques described herein. Such 25 representations, known as "IP cores," may be stored on a machine-readable physical medium and made available to various
18476
-11-
Customers or manufacturing facilities to load them onto manufacturing machines that make the logic or processor. It should be noted that many of the models shown here can of course be implemented using any suitable hardware and/or software languages (e.g., Java++, Swift, Objective-C, C,
QT, Perl, Python, JavaScript, etc.
<p dir="rtl">5 In some embodiments, one or more of the illustrative inventive computer-based devices of the present disclosure may include or be partially or completely included in at least one of a personal computer (PC), a laptop computer, a high-definition laptop computer, a tablet computer, a touch-screen tablet computer, a mobile computer, a handheld computer, a personal digital assistant (PDA), a cellular phone, a combined cellular phone/PDA, a television, a smart device (such as a smartphone or</p>
<p dir="rtl">10 (smart tablet or smart TV), mobile internet device (MID), receiver, data communication device, etc.</p>
As used herein, the term “server” should be understood to refer to a point of service that provides processing, database, and communications facilities. By way of example, and without limitation, the term “server” may refer to a single physical processor including communications, data storage device, and associated database facilities, or
<p dir="rtl">15 It can refer to a networked complex or complex of processors, network, and associated storage devices, as well as operating software, one or more database systems, and application software that supports the services provided by the server. Examples are cloud servers.</p>
In some embodiments, as described in detail herein, one or more of the illustrative inventive computer-based systems in accordance with the present disclosure may receive, process, and/or transmit and/or
<p dir="rtl">20 stores and/or transforms and/or produces and/or outputs any digital object and/or data unit (e.g., from within and/or outside of a particular application) that can have any suitable form, including, but not limited to, a file, contact, task, email, tweet, map, full application (e.g., calculator), etc. In some embodiments, as described in detail herein, one or more of the illustrative inventive computer-based systems according to the present disclosure can be implemented via</p>
<p dir="rtl">25 One or more different computer platforms such as, but not limited to: )1(</p>
18476
-12-
(4), Linux (3), NetBSD, OpenBSD, FreeBSD (2), AmigaOS, 4
(8), OS/2 (7), (Mac OS) OS 9(Solaris
(17), Embedded Linux (16), iOS (15), Firefox OS (14), BlackBerry OS, Windows Mobile (21), WebOS (20), Symbian, Tizen (19), Palm OS 5 (18).
Adobe (25), Adobe Flash (24), Adobe AIR (23), Windows Phone (22)
(BREW) Binary Runtime Environment for Wireless (26), Shockwave (27), JavaFX (31), Java Platforms (29), Cocoa Touch (28), (API) Cocoa
Mozilla Prism, XUL (34), Mono (33), Microsoft XNA (32), JavaFX Mobile Open Web (37), Silverlight (36), .NET Framework (35), and XULRunner 10 (41), SAP NetWeaver (40) (Qt (39), Oracle Database (38), Platform
Vexi (42), Smartface, and (43) Windows Runtime.
In some embodiments, the exemplary inventive computer-based systems and/or exemplary inventive computer-based devices according to the present disclosure may be configured to use wired circuitry.
<p dir="rtl">15 which may be used in lieu of, or in combination with, software instructions to implement features consistent with the disclosure principles. Accordingly, implementations consistent with the disclosure principles are not limited to any particular combination of hardware circuits and software. For example, many embodiments may be embodied in a variety of ways as a software component such as, but not limited to, a stand-alone software package, a combination of software packages, or a software package that is incorporated as a "tool" into a larger software product.</p>
<p dir="rtl">20 For example, illustrative software specifically programmed in accordance with one or more of the principles of the present disclosure may be downloadable from a network, e.g., a website, as a stand-alone product or as an add-on package for installation in an existing software application. For example, illustrative software specifically programmed in accordance with one or more of the principles of the present disclosure may also be available as a client-server software application, or as a software application<sup>^</sup>Be on the web. For example, you can</p>
18476
-13-
Also embodying specially programmed illustrative programs in accordance with one or more of the principles of the present disclosure as a software package installed on a device.
In some embodiments, the typical innovative computer-based systems/platforms, typical innovative computer-based devices and/or typical innovative computer-based components may be configured.
<p dir="rtl">5 Typical inventive computer according to the present disclosure for handling multiple concurrent users that may number, but are not limited to, at least 100 (for example, 100-999), at least 1,000 (for example, 1,000-9,999), at least 10,000 (for example, 10,000-99,999), at least 100,000 (for example, 100,000-999,999), at least 1,000,000 (for example, 1,0 ...).</p>
<p dir="rtl">10 Examples include, but are not limited to, 1000000-99999999, 10000000 or less (Example: 10000000-999999999), 100000000 or less (Example: 10000000-9999999999), 100000000 (Example: 100000000-1000000000).</p>
In some embodiments, the typical innovative computer-based systems and/or devices may be configured
<p dir="rtl">15 The typical inventive computer-based output according to the present disclosure is output to distinct and specially programmed graphical user interface implementations according to the present disclosure (e.g., desktop, web application, etc.). In many implementations of the present disclosure, the final output can be displayed on a display screen which may be, but is not limited to, a computer monitor, mobile device monitor, or the like. In various implementations, the display screen can be</p>
<p dir="rtl">20 The display is a holographic display. In various implementations, the display can be a transparent surface that may receive a visual projection. These projections can convey various forms of information, images, and/or objects. For example, these projections can be a visual overlay for a mobile augmented reality (MAR) application.</p>
As used herein, the terms “cloud,” “internet cloud,” “cloud computing,” “cloud architecture,” and similar terms are equivalent to at least one of the following: (1) a large number of devices,
18476
-14-
(2) Providing the ability to run a program or application on multiple connected computers (e.g., physical machines, virtual machines (VMs)) at the same time; (3) Network-based services, which appear to be provided by real server machines, but are in fact provided by
<p dir="rtl">5 Virtual machines (such as virtual servers), which are simulated by software running on one or more real machines (e.g., they are allowed to move around and quickly scale up (or down) without affecting the end user).</p>
In some embodiments, the exemplary inventive computer-based systems and/or exemplary inventive computer-based devices according to the present disclosure may be configured to securely store and/or transmit data 10 by using one or more encryption techniques (e.g., a double key
Private/Public, Triple Data Encryption Standard (3DES), Block Cipher Algorithms (eg,
CAST, RC5, RC2, IDEA and Skipjack), cryptographic hash algorithms (eg,
WHIRLPOOL (TTH), Tiger, SHA-2, SHA-1, RTR0, RIPEMD-160, MD5 (RNGs).
<p dir="rtl">15 The above examples are, of course, illustrative and not exhaustive.</p>
As used herein, the term “User” shall have the meaning of at least one user. In some embodiments, the terms “User,” “Subscriber,” “Consumer” or “Customer” shall be understood to refer to a user of an Application or Applications as described herein and/or a consumer of the data provided by a Data Provider. By way of example and not as a limitation, the terms “User” or “Subscriber” may refer to
<p dir="rtl">20 To a person who receives data provided by a data or service provider over the Internet in a browser session, or they may refer to an automated software application that receives data and stores or processes data.</p>
Figure 1 illustrates a block diagram of an illustrative computer-based system 100 for reducing bias in machine learning according to one or more of the embodiments of the present disclosure. However, not all of these components may be required to exercise one or more of the embodiments of the present disclosure, and variations in the arrangement and type of components may be made without
18476
-15-
Deviating from the scope or intent of various embodiments of the present disclosure. In some embodiments, the illustrative inventive computing devices and/or inventive computing components of the illustrative computer-based system 100 can be configured to manage a plurality of concurrent members and/or transactions, as detailed herein. In some embodiments, the computer system/platform can be based on
<p dir="rtl">5 Illustration 100 on a scalable computer and/or network architecture that includes various strategies for data evaluation, caching, searching, and/or database connection aggregation, including dynamic anomalous bias reduction (DOBR) as described in embodiments herein. An example of a scalable architecture is an architecture capable of running multiple servers.</p>
In some embodiments, referring to Figure 1, members 102-104 may include (for example)
<p dir="rtl">10 Example, clients (of the illustrative computer-based system 100) are virtually any computing device capable of receiving and sending a message over a network (e.g., a cloud network), such as the network 105, to and from another computing device, such as the servers 106 and 107, each other, and the like. In some embodiments, member devices 102-104 may be personal computers, multiprocessor systems, microprocessor-based or programmable consumer electronics,</p>
<p dir="rtl">15 A networked computer, and the like. In some embodiments, one or more member devices within member devices 102-104 may include computing devices that typically communicate using a wireless communications medium such as mobile phones, smartphones, pagers, walkie-talkies, radio frequency (RF) devices, infrared (CBs), integrated devices that combine one or more of the foregoing, or virtually any portable computing device,</p>
<p dir="rtl">20 And the like. In some embodiments, one or more of the member devices 102-104 may be a device capable of communicating using a wired or wireless communication medium such as a personal digital assistant, pocket PC, wearable computer, laptop, tablet, desktop computer, internet computer, video game console, pager, smartphone, ultra portable personal computer (UMPC), and/or any other device equipped to communicate via</p>
<p dir="rtl">25 Wired and/or wireless communication medium (e.g. 4G, 3G, NBIOT, RFID, NFC,</p>
18476
-16-
CDMA, WiMax, WiFi, GPRS, GSM, 5G, satellite, ZigBee, etc.). In some embodiments, one or more member devices within member devices 102-104 may include one or more applications, such as Internet browsers, mobile applications, voice calls, video games, video conferencing, and email, among others. In some embodiments, one or more member devices within member devices 102-104 may be configured to receive and transmit web pages and the like. In some embodiments, an exemplary browser application specially programmed in accordance with the present disclosure may be configured to receive and display graphics, text, multimedia, and the like, using virtually any web-based language, including, but not limited to, the Standard Generalized Markup Language (SMGL), such as HyperText Markup
<p dir="rtl">10 Language (HTML, Wireless Application Protocol (WAP), Mobile Markup Language</p>
(HDML), such as Wireless Markup Language (WMLScript, XML, JavaScript, WMLScript, and the like. In some embodiments, a member device within the member devices 102-104 may be specifically programmed by C++, C, QT, Net. Java, and/or other suitable programming language. In some embodiments, one or more member devices within member devices 102-104 may be specifically programmed 15 to include or implement a particular application to perform a variety of possible tasks, including but not limited to, the function of sending, browsing, searching, playing, streaming, or displaying various forms of content, including messages, images, video, and/or games stored locally or loaded locally.
In some embodiments, the demonstration network 105 may provide network access, data transmission and/or other services to any computing device associated with it. In some embodiments, the demonstration network 105 may include and implement at least one specialized network architecture that may be based in part on one or more standards developed by, but not limited to, the Global System for Mobile Communications (GSM) associations, the Internet Engineering Task Force (IETF), and the World Wide Web Interoperability Forum (WiMAX). In some embodiments, the demonstration network 105 may implement one or more of the GSM architectures, the General Packet Radio Service (GPRS) architecture, 25 the Universal Mobile Telecommunications System (UMTS) architecture, and the UMTS evolution referred to as the Evolution
18476
-17-
Long Term Evolution (LTE). In some embodiments, the demonstration network 105 may include and implement, instead of or in combination with one or more of the above elements, a WiMAX architecture defined by the WiMAX Forum. In some embodiments, and optionally, in combination with any embodiment described above or below, the demonstration network 105 may also include, for example, at least one area network
<p dir="rtl">5 Local area network (LAN), wide area network (WAN), Internet, virtual LAN (VLAN), enterprise LAN, Layer 3 virtual private network (VPN), enterprise IP network, or any combination thereof. In some embodiments, and optionally, in combination with any embodiment described above or below, at least one computer network connection may be transmitted over the illustrative network 105 based at least in part on one of the other communication modes, including but not limited to: NFC,</p>
<p dir="rtl">10 RFID, Narrowband Internet of Things (NBIOT, ZigBee, 3G, 4G, 5G, GSM), WiMax, WiFi, GPRS, satellite, and any combination thereof. In some embodiments, the illustrative network 105 may also include a mass storage device, such as a network attached storage device (NAS), storage area network (SAN), content delivery network (CDN), or other forms of machine- or computer-readable media.</p>
<p dir="rtl">15 In some embodiments, the demonstration server 106 or demonstration server 107 may be a web server (or series of servers) running a network operating system, examples of which may include, but are not limited to, Microsoft Windows Server, Novell NetWare, or Linux. In some embodiments, the demonstration server 106 or demonstration server 107 may be used for and/or to provide cloud and/or network computing. Although not illustrated in Figure 1, in</p>
<p dir="rtl">20 In some embodiments, the demonstration server 106 or the demonstration server 107 may include connections to external systems such as email, SMS, text messaging, advertising content providers, etc. Any of the features of the demonstration server 106 may also be implemented in the demonstration server 107 and vice versa.</p>
In some embodiments, one or more of the demonstration servers 106 and 107 may be specifically programmed 25 to perform, in a non-limiting example, as authentication servers, search servers, email servers,
18476
-18-
Social networking services servers, SMS servers, instant messaging servers, MMS servers, exchange servers, photo sharing services servers, advertising serving servers, financial/banking services servers, travel services servers or any similar suitable service base servers for the users of computers members 101-104.
<p dir="rtl">5 In some embodiments, and optionally, in combination with any embodiment described above or below, for example, one or more of the illustrative computer organ devices 102-104, the illustrative server 106 and/or the illustrative server 107 may include a specially programmed software module that can be configured to send, process and receive information using a scripting language, remote procedure call, email, tweet, short message service (SMS), multimedia messaging service</p>
<p dir="rtl">10 Multimedia (MMS), Instant Messaging (IM), Internet Relay Chat (IRC, mIRC, Jabber), API, Simple Object Access Protocol (SOAP), Shared Object Request Broker Architecture (HTTP), CORBA (Hypertext Transfer Protocol), REST (Representational State Transfer) or any combination thereof.</p>
Figure 2 illustrates a block diagram of another illustrative computer-based system/system 200 according to
<p dir="rtl">15 For one or more embodiments of the present disclosure. However, not all of these components may be required to practice one or more embodiments, and variations in the arrangement and type of components may be made without departing from the intent or scope of the various embodiments of the present disclosure. In some embodiments, the computing device 202a, 202b to 202n as described below comprise a computer-readable medium, such as random access memory (RAM) coupled to a processor 210 or FLASH memory. In some</p>
<p dir="rtl">20 Embodiments, the processor 210 may execute computer-executable program instructions stored in memory 208. In some embodiments, the processor 210 may include a microprocessor and/or an ASIC and/or a state machine. In some embodiments, the processor 210 may include, or may be in contact with, media, such as computer-readable media, which stores instructions that, when executed by the processor 210, can cause the processor 210 to execute a step or</p>
<p dir="rtl">25 More than the steps described here. In some embodiments, there are examples of readable media.</p>
18476
-19-
By computer includes, but is not limited to, another electronic, optical, or magnetic storage or transmission device capable of providing a processor, such as the processor 210 of the client 202a, with computer-readable instructions. In some embodiments, other examples of suitable media include, but are not limited to, a floppy disk, CD, DVD, magnetic disk,
<p dir="rtl">5 A memory chip, ASIC, RAM, ROM, a formatted processor, all optical media, all magnetic tape or other magnetic media, or any other medium through which a computer processor can read instructions. In addition, various other forms of computer-readable media can transmit or transfer instructions to a computer, including a router, private or public network, or other device or transmission channel, wired or wireless. In some embodiments,</p>
<p dir="rtl">10 Instructions can include code from any computer programming language, including, for example, Perl, Python, Java, Visual Basic, C++, JavaScript, etc.</p>
In some embodiments, the computing device 202a to 202n may also include a number of external or internal devices such as a mouse, CD-ROM, DVD, physical keyboard or 15 keyboard, display or other input or output devices. In some embodiments, examples of
The computing devices 202a through 202n (e.g., clients) may be any type of processor-based platform connected to the network 206 such as, but not limited to, personal computers, digital assistants, personal digital assistants, smartphones, pagers, digital tablets, laptop computers,
<p dir="rtl">20 Internet devices, and other processor-based devices. In some embodiments, the computing devices 202a through 202n can be specifically programmed using one or more application programs in accordance with one or more of the principles/methodologies detailed herein. In some embodiments, the computing devices 202a through 202n can operate on any operating system capable of supporting a browser or browser-enabled application, such as Windows™, Microsoft™, and/or Linux. In</p>
<p dir="rtl">25 Some embodiments, the computer organ devices 202A to 202N shown, may contain:</p>
18476
-20-
For example, personal computers running a browser application software such as Microsoft Corporation's Internet Explorer™, Apple Inc.'s Safari™, Mozilla Firefox, and/or Opera. In some embodiments, through the client computers 202a to 202n, users 212a to 212n may communicate via the network 206 5 to each other and/or to other systems and/or devices associated with the network 206. As shown in
Figure 2, illustrative server devices 204 and 213 may also be compared to the network 206. In some embodiments, one or more member computers 202a through 202n may be mobile clients.
In some embodiments, the one or more databases of illustrations 10 207 and 215 can be any type of database, including a database managed by a management system.
Databases (DBMS). In some embodiments, the DBMS-managed declarative database may be specifically programmed as an engine that controls the organization, storage, management, and/or retrieval of data in the database in question. In some embodiments, the DBMS-managed declarative database may be specifically programmed to provide the ability to query, backup, replicate, enforce rules, provide security, calculate, perform change and access logging, and/or automate
Optimization. In some embodiments, the DBMS-managed illustrative database can be selected from a FileMaker, Adaptive Server Enterprise, IBM DB2, Oracle, PostgreSQL, MySQL, Microsoft SQL Server, Microsoft Access, and NoSQL implementation. In some embodiments, the DBMS-managed illustrative database can be specifically programmed to select each relevant schema for each database in the DBMS, according to the model
A particular database of the present disclosure which may include a hierarchical model, a network model, a relational model, an object model, or some other suitable organization that may result in one or more applicable data structures which may include fields, records, files, and/or objects. In some embodiments, the declarative database managed by the DBMS may be specifically programmed to include 25 metadata about the stored data.
18476
-21-
In some embodiments, the illustrative computer-based systems/platforms, illustrative inventive computer-based devices, and/or illustrative inventive computer-based components according to the present disclosure may be specifically configured to operate in a cloud computing/architecture such as, but not limited to: Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and/or Software as a Service (SaaS). Figures 3 and 4 illustrate implementation architectures.
Illustrative architecture(s)/cloud computing in which the illustrative inventive computer-based platforms/systems, illustrative inventive computer-based devices, and/or illustrative inventive computer-based components, and/or illustrative inventive computer-based components according to the present disclosure can be specifically configured for the work.
<p dir="rtl">10 In innovative computer-based explanatory models of systems and/or devices, dynamic anomalous bias reduction (DOBR) can be used to improve the accuracy and understanding of general anomalous models specifically for benchmarking studies. However, it is a method that can be applied to a variety of analytical models where there is one or more independent variables and one dependent variable. The present disclosure and the models therein demonstrate the innovative application of DOBR to improve the accuracy of machine learning model predictions.</p>
<p dir="rtl">15 In models, DOBR is not a predictive model. Instead, in models, it is an additional method to predictive or explanatory models that can improve the accuracy of model predictions. In models, outliers identified by DOBR are based on the difference between the target variable supplied with data and the model-generated value. In order to identify outliers, through a pre-defined selection criterion, data records based on the outlier and dependent variables produced by the model are removed.</p>
<p dir="rtl">20 Further analysis can continue to permanently remove these records. However, in other embodiments of the inventive system and method, at each iteration of the model, the outlier identification process includes the entire data set such that all records are subjected to outlier scrutiny using the predictive model of the last iteration as determined by its calculation parameters. Accordingly, the explanatory models of the present invention reduce bias in the machine learning model, for example, by including</p>
<p dir="rtl">25 This is a complete dataset at each iteration to reduce the spread of selection bias to the training data. Hence,</p>
18476
-22-
Machine learning models can be trained and executed more accurately and efficiently to improve the operation of machine learning systems.
Figure 5 illustrates a framework diagram of an explanatory bias reduction system in machine learning according to one or more embodiments of the present disclosure.
<p dir="rtl">5 In some embodiments, the bias reduction system 300 may include a dynamic anomaly bias reduction (DOBR) component in the datasets being analyzed, e.g., machine learning engines. In some embodiments, DOBR provides an iterative process for removing anomalous records subject to a predetermined criterion. This condition represents a user-defined error acceptance value expressed as a percentage. This condition indicates the amount of error that the user is willing to accept in the model based on</p>
<p dir="rtl">10 Their vision and other analysis results will be described later in this discussion. A value of 100% indicates that all errors are accepted and no records will be removed in the DOBR process. If 0% is chosen, all records will be deleted. In general, error acceptance values in the range of 80 to 95% have been observed for industrial applications.</p>
In some embodiments, the user may interact with the bias reduction system 300 to manage the acceptance value.
<p dir="rtl">15 error via a user input device 308 and displaying the results via a display 312, among other user interaction behaviors using the display 312 and the user input device 308. Based on the error acceptance value, the bias reduction system 300 may analyze a data set 311 received in a database 310 or other storage unit in contact with the bias reduction system 300. The bias reduction system 300 may receive the data set 311 via a database 310 or device</p>
<p dir="rtl">20 Store and make predictions using one or more machine learning models with dynamic bias reduction to improve accuracy and efficiency.</p>
In some embodiments, the bias reduction system 300 includes a combination of hardware and software components, including, for example, storage and memory devices, caches, buffers, a bus, input/output (I/O) interfaces, processors, controllers, networking and communications devices,
18476
-23-
An operating system, kernel, device drivers, among other components. In some embodiments, the processor 307 is in communication with multiple other components to perform the functions of the other components. In some embodiments, each component has time scheduled on the processor 307 to perform the functions of the component, but in some embodiments, each component is scheduled to one or more processors in the processing system of the processor 307. In
<p dir="rtl">5 Other models, each component includes its own built-in processor.</p>
In some embodiments, the components of the bias reduction system 300 may include, for example, a DOBR motor 301 in communication with a model index 302 and a model library 303, a slope variable parameter library 305, a classifier parameter library 304 and a DOBR filter 306, among other possible components. Each component may include a combination of hardware and software to implement the functions of the components,
<p dir="rtl">10 Such as, memory and storage devices, processing devices, communication devices, input/output (I/O) interfaces, control units, networking and communication devices, an operating system, a kernel, device drivers, a set of instructions, among other components.</p>
In some embodiments, the DOBR engine 301 includes a model engine for instantiating and executing machine learning models. The DOBR engine 301 may access models for instantiation in the model library 303 15 by using the model index 302. For example, the model library may contain
303 A library of machine learning models that can be selectively accessed and instantiated for use by an engine such as the DOBR engine 301. In some embodiments, the model library 303 may include machine learning models, e.g., a support vector machine (SVM), a linear regression variable, a Lasso model, a decision tree regression variable, decision tree classifiers,
<p dir="rtl">20 Random Forest Regression, Random Forest Classifiers, K-Neighborhood Regression Variables, Classifiers</p>
Neighborhood K, gradient boosting gradients, gradient boosting classifiers, among other possible classifiers and gradients. For example, the model library 303 could import models according to the following example code 1:
fake code 1
18476
-24-
import sys
sys.path.append("analytics-lanxess-logic")
import numpy as np
import pandas as pd
import random, time 5
import xgboost as xgb
from xgboost import XGBClassifier,XGBRegressor
from scipy import stats
from scipy.stats import mannwhitneyu,wilcoxon
from sklearn.metrics import 10
mean_squared_error,roc_auc_score,classification_report,confusion_matrix
from sklearn import svm
from sklearn.svm import SVR, SVC
from sklearn.model_selection import train_test_split 15
from sklearn.linear_model import LinearRegression, Lasso
from sklearn.tree import DecisionTreeRegressor, DecisionTreeClassifier
from sklearn.ensemble import RandomForestRegressor,
RandomForestClassifier,BaggingClassifier,BaggingRegressor
18476
-25-
from sklearn.neighbors import KNeighborsRegressor ,
KNeighborsClassifier
from sklearn.ensemble import
GradientBoostingRegressor, GradientBoostingClassifier
5
from optimizers.hyperparameters.hyperband_optimizer import
Hyperband,HyperparameterOptimizer
from optimizers.hyperparameters.base_optimizer import
HyperparameterOptimizer
import warnings 10
from warnings import simplefilter
simplefilter(action='ignore', category=FutureWarning)
simplefilter(action='ignore', category=DeprecationWarning)
warnings.filterwarnings(module='numpy*' , action='ignore', category=DeprecationWarning) 15
warnings.filterwarnings(module='numpy*' , action='ignore', category=FutureWarning)
warnings.filterwarnings(module='scipy*', action='ignore', category=FutureWarning)
warnings.filterwarnings(module='scipy*' , action='ignore', 20 category=DeprecationWarning)
18476
-26-
warnings.filterwarnings(module='sklearn*', action='ignore',
category=DeprecationWarning)
However, in some embodiments, to facilitate access to the machine learning model library in the model library 303, the DOBR engine 301 may use a model index 302 that indexes each model with an identifier
<p dir="rtl">5 A model to be used as a function by the DOBR 301 engine. For example, models that include, for example, linear regression, XGBoost regression, support vector regression, Lasso, K-neighbor regression, bagging regression, gradient boosting regression, random forest regression, decision tree regression, among other regression and classification models, can be indexed by a number identifier and labeled with a specific name. For example, the code below shows an example of a model index code.</p>
<p dir="rtl">10 For use by Form 302 Index.</p>
fake code 2
model0 = LinearRegression()
model1 = xgb.XGBRegressor()
model2 = SVR()
model3 = Lasso() 15
model4 = KNeighborsRegressor()
model5 = BaggingRegressor()
model6 = GradientBoostingRegressor()
model7 = RandomForestRegressor()
model8 = DecisionTreeRegressor() 20
#
18476
-27-
ModelName0 = "Linear Regression"
ModelName1 = "XGBoost Regression"
ModelName2 = "Support Vector Regression"
ModelName3 = "Lasso"
ModelName4 = "K Neighbors Regression" 5
ModelName5 = "Bagging Regression"
ModelName6 = "Gradient Boosting Regression"
ModelName7 = "Random Forest Regression"
ModelName8 = "Decision Tree Regression"
<p dir="rtl">10 Other embodiments of the pseudocode of the model library 303 and the model index 302 are considered. In some embodiments, the program instructions are stored in the memory of the relevant model library 303 or the model index 302 and are temporarily stored in a cache for availability to the processor 307. In some embodiments, the DOBR engine 301 may use the model index 302 by accessing or calling the index via communications and/or input/output devices, where the index is used to call the models as functions.</p>
<p dir="rtl">15 From the Model 303 library via communications and/or input/output devices.</p>
In some embodiments, to facilitate optimization and customization of models called DOBR drive 301, the bias reduction system 300 may record model parameters in, for example, memory or storage, for example, hard disk drives, solid state disk drives, random access memory (RAM), flash storage, among other storage devices.
<p dir="rtl">20 and other memory. For example, the slope variable parameters may be recorded and modified in a slope variable parameter library 305. The slope variable parameter library 305 may then include storage and communications devices configured with sufficient memory and bandwidth to store, modify, and communicate a number of</p>
18476
-28-
A large number of parameters for multiple regression variables, for example, in real time. For example, for each regression machine learning model instantiated by the DOBR engine 301, its parameters may be initialized and updated in the regression variable parameter library 305. In some embodiments, the user may, via the user input device 308, create an initial set of parameters.
<p dir="rtl">5 However, in some embodiments, the initial set of parameters may be pre-defined or randomly generated. When creating an instance of a regression machine learning model, the DOBR engine 301 may associate a model, as identified in the model index 302 , with a set of parameters in the regression variable parameter library 305 . For example, the DOBR engine 301 may call a set of parameters based on, for example, an identification number (ID) associated with a particular regression model.</p>
<p dir="rtl">10 For example, the regression parameter library 305 could define the parameters of each regression model similar to the pseudocode 3 below:</p>
fake code 3
#from utilities.defaults import DefaultParameters
#print(DefaultParameters(ctr=0).__dict__)
#!conda install -y -c conda-forge xgboost 15
def gen_params(id):
#XGBoost
if id==1:
"" default parameters - best achieved in prototyping XGBOOST ""
HYPERPARAMETERS = {"objective": "reg:linear", 20
"tree_method": "exact",
"eval_metric": "rmse",
18476
-29-
eta": 1
"gamma": 5,
"max_depth": 2,
"colsample_bytree": .5,
"colsample_bylevel": .5,
"min_child_weight": 1,
"subsample": 1,
"reg_lambda": 1,
"reg_alpha": 0,
"silent": 1}
""" fixed parameters which will not change in optimization """ FIXED = {"objective": "reg:linear", "tree_method": "exact", "eval_metric": "rmse"}
""boundaries & types of optimisable parameters"" BOUNDARIES = {"eta": (0, 1, np.float64), "gamma": (0, 100, np.float64),
10
15
18476
-30-
"max_depth": (1, 30, np.int32),
"colsample_bytree": (0, 1, np.float64), "colsample_bylevel": (0, 1, np.float64), "min_child_weight": (0, 100, np.int32),
"subsample": (0, 1, np.float64), 5
"reg_lambda": (0, 1, np.float64), "reg_alpha": (0, 1, np.float64)} elif id==2:
#SVR
""default parameters -"" 10
HYPERPARAMETERS = {"kernel": "rbf",
"cache_size": 100000,
"C": 0.5, "gamma": 0.023 }
15
""" fixed parameters which will not change in optimization """ FIXED = {"kernel": "rbf", "cache_size": 100000,
"tol": 0.00001 }
18476
-31-
""boundaries & types of optimizable parameters""
BOUNDARIES = { "C": (0.01, 1000, np.float64),
"gamma": (0.001, 100, np.float64)}
# "epsilon": (0.001, 100, np.float64) 5
elif id==3:
# LASSO
""" default parameters - """
HYPERPARAMETERS = {"fit_intercept": "False", 10
"max_iter": 100000,
"tol": 0.0001,
"alpha": 25}
""fixed parameters which will not change in optimization"" 15
FIXED = {"fit_intercept": "False",
"max_iter": 100000,
"tol": 0.0001 }
""boundaries & types of optimizable parameters""
18476
-32-
BOUNDARIES = {"alpha": (0.1, 100, np.float64) }
elif id==4:
# KNN PARAMETERS 5
""" default parameters - """
HYPERPARAMETERS = { "algorithm": "auto",
"n_neighbors": 7,
"leaf_size": 30}
10
""" fixed parameters which will not change in optimization """ FIXED = {"algorithm": "auto"}
""" boundaries & types of optimisable parameters """ BOUNDARIES = {"n_neighbors": (3 , 51, np.int32), 15 "leaf_size": (2 , 500, np.int32)}
elif id==5:
# Bagging Regression
18476
-33-
HYPERPARAMETERS = { "bootstrap_features": "False",
"bootstrap": "True",
"n_estimators": 21,
"max_samples": 23}
5
""" fixed parameters which will not change in optimization """ FIXED = { "bootstrap_features": "False",
"bootstrap": "True"}
"" boundaries & types of optimizable parameters "" 10
BOUNDARIES = {"n_estimators": (1, 50, np.int32),
"max_samples": (1, 50, np.int32)}
elif id==6:
#GRADIENT BOOSTING PARAMETERS 15
""" default parameters - """
HYPERPARAMETERS = {"criterion": "friedman_mse",
"min_impurity_split": 1.0e-07,
"max_features": "auto",
18476
-34-
"learning_rate": 0.2,
"n_estimators": 100,
"max_depth": 10}
""fixed parameters which will not change in optimization"" 5
FIXED = {"criterion": "friedman_mse",
"min_impurity_split": 1.0e-07, "max_features": "auto"}
"" boundaries & types of optimizable parameters "" 10
BOUNDARIES = {"learning_rate": (0.01, 1, np.float64),
"n_estimators": (50, 500, np.int32),
"max_depth": (1, 50, np.int32)}
elif id==7:
# RANDOM FOREST PARAMETERS 15
""" default parameters - """
HYPERPARAMETERS = {"bootstrap": "True",
"criterion": "mse",
"n_estimators": 100,
18476
-35-
"max_features": 'auto',
"max_depth": 50, "min_samples_leaf": 1, "min_samples_split": 2}
5
""" fixed parameters which will not change in optimization """ FIXED = {"bootstrap": "True", "criterion": "mse",
"max_features": 'auto' }
10
""boundaries & types of optimizable parameters""
BOUNDARIES = {"n_estimators": (1, 1000, np.int32),
"max_depth": (1, 500, np.int32),
"min_samples_leaf": (1, 50, np.int32),
"min_samples_split": (2, 50, np.int32)} 15
else:
<p># DECISION TREE PARAMETERS</p>
""" default parameters - """
18476
-36-
HYPERPARAMETERS = {"criterion": "mse",
"max_features": "auto",
"max_depth": 2,
"min_sample_leaf ": 0.25,
"min_samples_split": 2 } 5
""fixed parameters which will not change in optimization""
FIXED = {"criterion": "mse",
"max_features": "auto"}
10
""boundaries & types of optimizable parameters""
BOUNDARIES = { "max_depth": (1, 500, np.int32),
"min_samples_leaf": (1, 50, np.int32),
"min_samples_split": (2, 50, np.int32)}
15
return HYPERPARAMETERS,FIXED,BOUNDARIES
Similarly, in some embodiments, classifier parameters may be recorded and modified in classifier parameter library 304. Accordingly, classifier parameter library 304 may contain storage and communications devices configured with sufficient memory and bandwidth to store, modify, and communicate a large number of regression variable parameters.
<p dir="rtl">20 Multiple, for example, in real time. For example, for each machine learning model</p>
18476
-37-
For classification, an instance of the DOBR engine 301 is created, the parameters of which may be initialized and updated in the regression variable parameter library 305. In some embodiments, the user may, via the user input device 308, create an initial set of parameters. However, in some embodiments, the initial set of parameters may be predefined. When an instance of a regression machine learning model is created, the user may
<p dir="rtl">5 DOBR engine 301 associates a model, as specified in model index 302, with a set of parameters in the regression variable parameter library 305. For example, DOBR engine 301 may call a set of parameters based on, for example, an identification number (ID) associated with a particular regression model. For example, the regression variable parameter library 305 may specify parameters for each regression model similar to the spoof code 4 below:</p>
10 Fake Code 4
def gen_paramsClass(II):
<p># XGBoost CLASSIFER PARAMETERS</p>
if II==0:
""default parameters - best achieved in prototyping""
HYPERPARAMETERS = {"objective": "binary:hinge", 15
"tree_method": "exact",
"eval_metric": "error",
"n_estimators": 5,
"eta": 0.3,
"gamma": 0.1, 20
"max_depth": 5,
18476
-38-
"min_child_weight": 5,
"subsample": 0.5,
"scale_pos_weight": 1,
"silent": 1}
""fixed parameters which will not change in optimization""5
FIXED = { "objective": "binary:hinge",
"tree_method": "exact",
"eval_metric": "error"}
""boundaries & types of optimizable parameters""
BOUNDARIES = { "eta": (0, 10, np.float64), 10
"gamma": (0, 10, np.float64),
"min_child_weight": (0, 50, np.float64),
"subsample": (0, 1, np.float64),
"n_estimators": (1,1000, np.int32),
"max_depth": (1, 1000, np.int32), 15
"scale_pos_weight": (0, 1, np.float64) }
else:
<p># RANDOM FOREST CLASSIFIER PARAMETERS</p>
""" default parameters - """
18476
-39-
HYPERPARAMETERS = {"bootstrap": "True",
"n_estimators": 500,
"max_features": 'auto',
"max_depth": 200, 5
"min_samples_leaf": 1,
"min_samples_split": 2 }
""fixed parameters which will not change in optimization""
FIXED = {"bootstrap": "True",
"max_features": "auto" } 10
""boundaries & types of optimizable parameters""
BOUNDARIES = {"n_estimators": (10, 1000, np.int32),
"max_depth": (10, 50, np.int32),
"min_samples_leaf": (1, 40, np.int32), 15
"min_samples_split": (2, 40, np.int32)}#
return HYPERPARAMETERS,FIXED,BOUNDARIES
In some embodiments, by calling and receiving a set of models from the model library 303 via the model index 302 and relevant parameters from the regression variable parameter library 305 and/or the
<p dir="rtl">20 Classifier parameter 304, DOBR engine 301 can load one or more models.</p>
18476
-40-
The instance and the models that are initialized, for example, in the cache of the DOBR engine 301. In some embodiments, the data set 311 may then be loaded from the database 310 into, for example, the same cache or another storage device of the DOBR engine 301. The processor 307 or the processor in the DOBR engine 301 may then execute
<p dir="rtl">5 Each model transforms the dataset 311 into, for example, a prediction of activity-related data values that characterize the outcomes or parameters of an activity based on certain input attributes related to the activity. For example, energy usage of appliances in residential and/or commercial environments, compressive strength of concrete in a variety of applications and formats, object or image recognition, speech recognition, or other machine learning applications. For example, the DOBR engine 301 may</p>
<p dir="rtl">10 Modeling appliance energy usage based on a dataset 311 of historical energy usage, time of year, time of day, location, among other factors. The DOBR engine 301 may call a set of regression variables from the model library 303 via the model index 302 connected to a bus to the DOBR engine 301. The DOBR engine 301 may then call a parameter file or register associated with the regression variable to estimate appliance energy usage in the regression variable parameter library 305</p>
<p dir="rtl">15 Connected to the DOBR motor bus 301. The DOBR motor 301 may then use the processor 307 to predict future energy consumption based on models and model parameters, time and date, location or any other factor and combinations thereof.</p>
Similarly, for example, a 301 DOBR engine could be a concrete compressive strength model based on a 311 dataset of concrete materials, time of year, time of day,
<p dir="rtl">20 location, moisture, curing time, age, among other factors. The DOBR engine 301 may call a set of slope variables from the model library 303 via the model index 302 connected to a DOBR engine 301 bus. The DOBR engine 301 may then call a parameter record or file associated with the slope variables to estimate the compressive strength of concrete in the slope variable parameter library 305 connected to a DOBR engine 301 bus. The DOBR engine 301 may then use</p>
18476
-41-
Processor 307 to predict future concrete compressive strength based on models and model parameters for a specific concrete formulation, time and date, location, or other factor and combinations thereof.
In another example, the DOBR engine 301 can recognize speech based on a dataset 311 of actual field speech and text, among other factors. It can call
<p dir="rtl">5 DOBR engine 301 receives a set of classifiers from the model library 303 via the model index 302 connected to a bus of DOBR engine 301. DOBR engine 301 may then call a record or parameter file associated with the speech recognition classifiers in the classifier parameter library 304 connected to a bus of DOBR engine 301. DOBR engine 301 may then use processor 307 to predict transcriptions of the recorded speech data based on the models and model parameters for a single-word or</p>
<p dir="rtl">10 more.</p>
In another example, the DOBR engine 301 may automatically predict display settings for medical images based on a dataset 311 of settings for multiple display parameters across imaging and/or visualizations, among other factors, as described in U.S. Pat. 10,339,695, which is incorporated herein in its entirety by reference for all purposes. The DOBR engine 301 may call
<p dir="rtl">15 301 DOBR then a set of classifiers from the model library 303 via the model index.</p>
302 connected to a DOBR engine 301 bus. The DOBR engine 301 may then call a classifier-related parameter register or file to display settings in the classifier parameter library 304 connected to a DOBR engine 301 bus. The DOBR engine 301 may then use the processor 307 to predict display setting data based on the models and model parameters for a set of
<p dir="rtl">20 One or more medical data.</p>
In another example, the DOBR engine 301 may automatically control machines based on a data set 311 of machine control command results, simulated machine control command results, among other factors, as described in U.S. Pat. 10,317,854, which is incorporated herein in its entirety by reference for all purposes. The DOBR engine 301 may call upon a set of models
<p dir="rtl">25 Descend from the model library 303 via the model index 302 connected to a DOBR motor carrier.</p>
18476
-42-
<p dir="rtl">301. The DOBR engine may then call a parameter register or file associated with the regression model for automatic control in a regression variable parameter library 305 connected to a bus to the DOBR engine 301. The DOBR engine 301 may use the processor 307 to predict the success or failure of certain control commands based on models and model parameters for a set of control commands, environmental information, sensor data, and/or 5 command simulations.</p>
In some embodiments, the bias reduction system 300 may execute machine learning models in a cloud environment, for example, as a cloud service for remote users. Such cloud service may be designed to support large numbers of users and a wide range of algorithms and problem sizes, including those described above, as well as other potential models and datasets.
<p dir="rtl">10 and parameter tuning operations for a user use case, as described in U.S. Patent No. 10,452,992, are incorporated herein in their entirety by reference for all purposes. In one embodiment, a number of programming interfaces (e.g., APIs) may be specified by the service in which the bias reduction system 300 is implemented, which guide non-expert users to get started with machine learning best practices relatively quickly, without the users having to spend a lot of</p>
<p dir="rtl">15 time and effort on tuning models, or in learning advanced statistics or AI techniques. Interfaces can, for example, allow non-experts to rely on default settings or parameters for various aspects of the procedures used in building, training, and using machine learning models, where the default settings are derived from one or more sets of parameters in the classifier parameter library 304 and/or the regression variable parameter library 305 for similar models.</p>
<p dir="rtl">20 For the individual user. The default settings or parameters may be used as a starting point to customize a user's machine learning model using training with user datasets via the DOBR engine 301 and optimizer 306. At the same time, users may customize the parameters or settings that they wish to use for different types of machine learning tasks, such as input history processing, feature processing, model building, execution, and evaluation. In at least some embodiments, in addition to or</p>
<p dir="rtl">25 Instead of using pre-defined libraries that implement different types of machine learning tasks, in addition to</p>
18476
-43-
In addition, the Cloud Service Bias Reduction System 300 may include built-in extensible capabilities for the service, for example, by registering custom functions with the service. Depending on the business needs or objectives of the customers implementing such modules or custom functions, the modules may in some cases be shared with other users of the service, while in
<p dir="rtl">5 Other cases where custom modules may be limited to their implementers/owners.</p>
In some embodiments, whether implemented as a cloud service, on-premises or remote system, or in any other system architecture, the bias reduction system 300 may include models in the model library 303 that enable the training and execution of a consistent ensemble approach to a machine learning model, as described in U.S. Pat. 9,646,262, incorporated herein in its entirety by reference for all purposes. Such approach may be
<p dir="rtl">10 Useful for applications to data analytics using electronic datasets of data.</p>
Electronic activity. In some embodiments, the database 310 may contain one or more structured or unstructured data sources. The unsupervised learning module, in certain embodiments, is configured to aggregate an unstructured dataset into a structured dataset using a set of unsupervised learning techniques, for example, into a consistent set of
<p dir="rtl">15 Models from the Model Library 303. For example, an unsupervised learning module is configured to aggregate an unstructured dataset into multiple versions of a structured dataset, while a supervised learning module is configured, in certain models, to generate one or more consistent machine learning ensembles based on each of the multiple versions of a structured dataset and to select the consistent machine learning ensemble that exhibits the highest predictive performance according to,</p>
<p dir="rtl">20 For example, for a model error after training each model in each consistent set using the DOBR engine 301 and the optimizer 306.</p>
An example of the DOBR engine 301 instructions to control devices to make predictions based on data set 311 is illustrated in the pseudocode 5 shown below:
fake code 5
18476
-44-
filename = 'energydataBase'
filename = 'Concrete_Data'
path = '.'
filetype = '.csv'
path1 = filename + filetype 5
data = pd.read_csv(path1).values
YLength = len(data)
X_Data = data[:, 1:]
y_Data = data[:,0]
# 10
<p># ***** Set Run Parameters *****</p>
#
ErrCrit = 0.005
trials = 2
15
list_model = [ model0, model1, model2, model3, model4 ]
list_modelname = [ ModelName0, ModelName1, ModelName2,
ModelName3, ModelName4]
Acceptance = [87.5, 87.5, 87.5, 87.5, 87.5]
18476
-45-
#
mcnt = -1
for model in list_model:5
f = open("DOBR04trainvaltestRF"+".txt","a")
mcnt += 1
print("- ...
mcnt,list_modelname[mcnt])
timemodelstart = time. time()10
Error00 = [0]*trials
PIM = [0]*trials
modelfrac = [0]*trials
DOBRFULL0,DOBRFULL0a,DOBRFULL0e = ([0] * trials for i in range(3))15
DOBRFULL1,DOBRFULL1a,DOBRFULL1e = ([0] * trials for i in range(3))
DOBRFULL2,DOBRFULL2a,DOBRFULL2e = ([0] * trials for i in range(3))
#
<p># Bootstrapng Loop starts here</p>
18476
-46-
X_train, X_temp, y_train, y_temp = train_test_split(X_Data, y_Data, test_size = 0.60)
if mcnt > 0:
new_paramset = gen_params(mcnt) 5
hyperband = Hyperband(X_train, y_train, new_paramset[0], new_paramset[1], new_paramset[2])
hyperband.optimize(model)
<p># print("Best parameters", hyperband.best_parameters)</p>
RefModel = model.set_params(**hyperband.best_parameters) 10
else:
RefModel = model
print(RefModel,file=f)
#
for mc in range(0,trials): 15
x_val, x_test, y_val, y_test = train_test_split(X_temp, y_temp, test_size = 0.20)
timemodelstart1 = time.time()
len_yval = len(y_val)
len_ytest = len(y_test) 20
18476
-47-
Errmin = 999
Errvalue = 999
cnt = 0
#
BaseModel = RefModel.fit(x_val,y_val).predict(x_test) 5
Error00[mc] = (mean_squared_error(BaseModel, y_test))**0.5
#
DOBRModel = RefModel.fit(x_val,y_val).predict(x_val)
Errorval = (mean_squared_error(DOBRModel, y_val))**0.5
print("Train Error", Error00[mc],"Test Error", Errorval," Ratio: 10
",Error00[mc]/Errorval,"mc=",mc)
<p># Data_xin0_values = x_val</p>
<p># Data_yin0_values = y_val</p>
<p># XinBest = x_val</p>
<p># YinBest = y_val 15</p>
#
rmsrbf1 = Error00[mc]
while Errvalue > ErrCrit:
cnt += 1
18476
-48-
timemodelstart1 = time.time()
if cnt > 500:
print("Max iter. cnt for Error Acceptance: ",Errvalue,Acceptance[mcnt])
break
# 5
<p># Absolute Errors & DOBR Filter</p>
#
AError = RMS(DOBRModel, y_val)
inout1 = DOBR(AError, Acceptance[mcnt])
# 10
Data_yin_scrub, dumb1 = scrub1(inout1, y_val)
Data_xin_scrub, dumb2 = scrub2(inout1, x_val)
DOBR_yin_scrub, dumb3 = scrub1(inout1, DOBRModel)
rmsrbf2 = (mean_squared_error(DOBR_yin_scrub ,Data_yin_scrub) )**0.5 15
#
if rmsrbf2 < Errmin:
<p># XinBest = Data_xin0_values</p>
<p># YinBest = Data_yin0_values</p>
18476
-49-
Errmin = rmsrbf2
Errvalue = abs(rmsrbf2 - rmsrbf1)/rmsrbf2
<p># print(cnt,Errvalue," ",rmsrbf2,rmsrbf1,sum(inout1)/len_yval)</p>
rmsrbf1 = rmsrbf2
DOBRModel = 5
RefModel.fit(Data_xin_scrub,Data_yin_scrub).predict(x_val)#<---------- ------
<p># Data_xin0_values = Data_xin_scrub</p>
<p># Data_yin0_values = Data_yin_scrub</p>
# 10
<p>#DOBRModel =</p>
RefModel.fit(Data_xin_scrub,Data_yin_scrub).predict(x_val)
<p># AError = RMS(DOBRModel, y_val)</p>
<p># inout1 = DOBR(AError, Acceptance[mcnt])</p>
print( " Convergence in ",cnt," iterations with Error Value = ",Errvalue) 15
#
#++++++++++++++++++++++++++++++++++++++++++++++++++++++++ +++++++++++++++++++++++++++++++++++++++++++++++++
if mc == mc:
timemodelstart2 = time.time() 20
18476
-50-
new_paramset = gen_paramsClass(1)
hyperband = Hyperband(np.array(x_val), np.array(inout1), new_paramset[0], new_paramset[1], new_paramset[2])
modelClass = RandomForestClassifier() #xgb.XGBClassifier()
hyperband.optimize(modelClass, True) 5
Classmodel = modelClass.set_params(**hyperband.best_parameters)
print(hyperband.best_parameters,file = f)
print(hyperband.best_parameters)
#
inout2 = Classmodel.fit(x_val, inout1).predict(x_test) 10
modelfrac[mc] = sum(inout1)/len_yval
PIM[mc] = sum(inout2)/len_ytest
#
<p># MODEL DOBR CENSORED DATASETS</p>
# 15
Data_yin_scrub, Data_yout_scrub = scrub1 (inout1, y_val)
Data_xin_scrub, Data_xout_scrub = scrub2 (inout1, x_val)
#
<p># TEST DOBR CENSORED DATASET</p>
18476
-51-
Data_xtestin_scrub, Data_xtestout_scrub = scrub2 (inout2, x_test)
y_testin_scrub, y_testout_scrub = scrub1 (inout2, y_test)
y_test_scrub = [*y_testin_scrub, *y_testout_scrub]
#
<p># DOBR INFORMATION APPLIED BASE MODEL PREDICTOR 5</p>
DATASET
BaseModel_yin_scrub, BaseModel_yout_scrub = scrub1(inout2, BaseModel)
#
DOBR_Model_testin = model.fit(Data_xin_scrub, Data_yin_scrub 10).predict(Data_xtestin_scrub)
if len(y_test) == sum(inout2):
DOBR_Model0 = DOBR_Model_testin
DOBR_Model1 = DOBR_Model_testin
DOBR_Model2 = BaseModel_yin_scrub 15
print("inout2:",sum(inout2),"len = ",len(y_test))
else:
DOBR_Model_testout = model.fit(Data_xout_scrub, Data_yout_scrub).predict(Data_xtestout_scrub)
DOBR_Model0 = [*DOBR_Model_testin, *DOBR_Model_testout ] 20
18476
-52-
DOBR_Model1 = [*DOBR_Model_testin , *BaseModel_yout_scrub]
DOBR_Model2 = [*BaseModel_yin_scrub, *DOBR_Model_testout ]
#
DOBRFULL0[mc] = (mean_squared_error(DOBR_Model0, y_test_scrub))**0.5 5
DOBRFULL1[mc] = (mean_squared_error(DOBR_Model1, y_test_scrub))**0.5
DOBRFULL2[mc] = (mean_squared_error(DOBR_Model2, y_test_scrub))**0.5
# 10
ModelFrac = np.mean(modelfrac ,axis=0)
Error00a = np.mean(Error00 ,axis=0)
DOBRFULL0a = np.mean(DOBRFULL0 ,axis=0)
DOBRFULL1a = np.mean(DOBRFULL1 ,axis=0)
DOBRFULL2a = np.mean(DOBRFULL2, axis=0) 15
Error00e = 1.96 * stats.sem(Error00 ,axis=0)
DOBRFULL0e = 1.96 * stats.sem(DOBRFULL0 ,axis=0)
DOBRFULL1e = 1.96 * stats.sem(DOBRFULL1 ,axis=0)
DOBRFULL2e = 1.96 * stats.sem(DOBRFULL2 ,axis=0)
# 20
18476
-53-
PIM_Mean = np.mean(PIM)
PIM_CL = 1.96 * stats.sem(PIM)
#
print(" "+ list_modelname[mcnt], " # of Trials =",trials,file=f)
print(Classmodel,file=f) 5
print(" Test Dataset Results for {0:3.0%} of Data Included in DOBR
Model {1:3.0%} ± {2:4.1%} "
.format(ModelFrac,PIM_Mean, PIM_CL),file=f)
print(" Base Model={0:5.2f} ± {1:5.2f} DOBR_Model #1 = {2:5.2f} ± 10
{3:5.2f}"
.format(Error00a,Error00e,DOBRFULL0a, DOBRFULL0e),file=f)
print("DOBR_Model #2 = {0:5.2f} ± {1:5.2f}".format(DOBRFULL1a,
DOBRFULL1e),file=f)
print("DOBR_Model #3 = {0:5.2f} ± {1:5.2f}".format(DOBRFULL2a, 15
DOBRFULL2e),file=f)
print(" "+ list_modelname[mcnt], " # of Trials =",trials)
print(Classmodel,file=f)
print(" Test Dataset Results for {0:3.0%} of Data Included in DOBR
Model {1:3.0%}±{2:4.1%}" 20
18476
-54-
.format(ModelFrac,PIM_Mean, PIM_CL))
print(" Base Model={0:5.2f}± {1:5.2f} DOBR_Model #1 = {2:5.2f}±
{3:5.2f}"
.format(Error00a,Error00e,DOBRFULL0a,DOBRFULL0e))
print("DOBR_Model #2 = {0:5.2f} ± {1:5.2f}".format(DOBRFULL1a, 5
DOBRFULL1e))
print("DOBR_Model #3 = {0:5.2f} ± {1:5.2f}".format(DOBRFULL2a,
DOBRFULL2e))
print("++++++++++++++++++++++++++++++++++++++++++++++++++ ++++++ 10
++++++++++++++++++++++++++++++++++++++++++++++++++")
#
f.close()
modeltime = (time.time() - timemodelstart) / 60
print("Total Run Time for {0:3} iterations = {1:5.1f} 15
min".format(trials,modeltime))
However, in some models, outliers in the dataset 311 can reduce the accuracy of the implemented models, resulting in an increased number of training iterations. To improve accuracy and efficiency, the DOBR engine 301 can include a DOBR filter 301b to dynamically test for data point errors.
<p dir="rtl">20 in the data set to identify outliers. The outliers can then be removed to provide a more accurate or representative data set 311. In some embodiments, the DOBR filter 301 may provide a recursive mechanism for removing outlier data points subject to a predetermined criterion, for example, a fault tolerance value</p>
18476
-55-
The user-specified error acceptance value described above and provided, for example, by a user via the user input device 308. In some embodiments, the user-specified error acceptance value is expressed as a percentage where, for example, a value of 100% indicates that all errors will be accepted and no data points will be removed by the filtering factor 301b, while a value of
<p dir="rtl">5 For example, 0% to remove all data points. In some embodiments, filter 301B may be initialized with an error acceptance value in the range between, for example, about 80% and about 95%. For example, filter 301B may be initialized to perform functions as shown in spoof code 6 below:</p>
fake code 6
# Absolute Errors & DOBR Filter 10
#
AError = RMS(DOBRModel, y_val)
inout1 = DOBR(AError, Acceptance[mcnt])
#
Data_yin_scrub, dumb1 = scrub1(inout1, y_val) 15
Data_xin_scrub, dumb2 = scrub2(inout1, x_val)
DOBR_yin_scrub, dumb3 = scrub1(inout1, DOBRModel)
rmsrbf2 = (mean_squared_error(DOBR_yin_scrub ,Data_yin_scrub) )**0.5
# 20
if rmsrbf2 < Errmin:
18476
-56-
<p># XinBest = Data_xin0_values</p>
<p># YinBest = Data_yin0_values</p>
Errmin = rmsrbf2
Errvalue = abs(rmsrbf2 - rmsrbf1)/rmsrbf2
# print(cnt,Errvalue," ",rmsrbf2,rmsrbf1,sum(inout1)/len_yval) 5
rmsrbf1 = rmsrbf2
DOBRModel =
RefModel.fit(Data_xin_scrub,Data_yin_scrub).predict(x_val)#<---------- ------
<p># Data_xin0_values = Data_xin_scrub 10</p>
<p># Data_yin0_values = Data_yin_scrub</p>
#
<p>#DOBRModel =</p>
RefModel.fit(Data_xin_scrub,Data_yin_scrub).predict(x_val)
# AError = RMS(DOBRModel, y_val) 15
# inout1 = DOBR(AError, Acceptance[mcnt])
print( " Convergence in ",cnt," iterations with Error Value = ",Errvalue)
#
In some embodiments, DOBR filter 301 operates in conjunction with enhancer 306, which is
<p dir="rtl">20 configured to identify the error and optimize the parameters for each model in the regression parameter library 305 and the classifier parameter library 304. Then, in some embodiments, the optimizer 306 may identify the model and transfer</p>
18476
-57-
error to the filter 301b of the DOBR engine 301. Thus, in some embodiments, the optimizer 306 may, for example, have storage and/or memory devices and communications devices with sufficient memory capacity and bandwidth to receive the dataset 311 and the model predictions and determine, for example, outliers, convergence, error, absolute value error, among other error measures.
<p dir="rtl">5 Other. For example, the function optimizer 306 can be configured as shown in the spoof code 7 below:</p>
fake code 7
def DOBR(AErrors,Accept):
length = len(AErrors)
Inout = [1]*length 10
AThres = stats.scoreatpercentile(AErrors, Accept)
for i in range(0,length):
if AErrors[i] > AThres:
Inout[i]=0
return Inout 15
def RMS(Array1,Array2):
length = len(Array1)
Array3 = [0 for m in range(0,length)]
for i in range(0,length):
Array3[i] = (Array1[i] - Array2[i])**2 20
18476
-58-
return Array3
def scrub1(IO,ydata):
lendata = len(ydata)
outlen = sum(IO)
Yin = []*output 5
Yout = []*(lendata - output)
for i in range(0,lendata):
if IO[i] > 0:
Yin.append(ydata[i]) else: 10
Yout.append(ydata[i])
return Yin,Yout
def scrub2(IO,Xdata):
lendata = len(Xdata)
inlen = sum(IO) 15
outlen = len(IO) - inlen
cols = len(Xdata[0])
Xin = [[0 for k in range(cols)] for m in range(inlen )]
Xout = [[0 for k in range(cols)] for m in range(outlen)]
18476
-59-
irow = -1
jrow = -1
for i in range(0,lendata):
if IO[i] > 0:
irow += 1 5
for j in range(0,cols):
Xin[irow][j] = Xdata[i][j]
else:
jrow += 1
for k in range(0,cols): 10
Xout[jrow][k] = Xdata[i][k]
return Xin,Xout
In some embodiments, the bias reduction system 300 may then return to the user via, for example,
Example, Display 312, Machine Learning Model Predictions, Outlier Analysis, Convergence of Predictions, From
<p dir="rtl">15 Among other data that the 301 DOBR engine produces in a more accurate and efficient way due to the reduction in</p>
Outliers that bias predictions.
Figure 6 illustrates a flowchart of an illustrative inventive methodology according to one or more embodiments of the present disclosure.
The DOBR, like the 301 DOBR engine and 301B filter described above, provides a repetitive process.
<p dir="rtl">20 To remove anomalous records subject to a pre-defined criterion. This condition is a false acceptance value.</p>
User-defined expressed as a percentage. This condition indicates the amount of error that is desired.
18476
-60-
The user accepts the model based on their insights and other analysis results that will be described later in this discussion. A value of 100% indicates that all errors are accepted and no records will be removed in the DOBR process. If zero is chosen, all records will be removed. In general, error acceptance values in the range of 80 to 95% have been observed for industrial applications.
<p dir="rtl">5 However, in some models, it should also be noted that if the data set contains no outliers, DOBR does not provide any value. But in practical situations it is rare for an analyst to have this knowledge before working with a data set. As will be explained later in this discussion, methodology models can also determine what percentage of the data set represents outliers for the model. This pre-analysis step can help determine the appropriate value for accepting an error of 10 or if there are outliers at all.</p>
The following steps demonstrate the basic DOBR method as applied to a complete dataset.
Pre-analysis: In one model, we first choose an error acceptance criterion, suppose we choose KM = 80%0 (how to determine this value from the data will be explained after explaining the DOBR method) and then the error acceptance criterion, (0)0 is determined according to, for example, Equation 1 below:
15 The equation [(-44 4.2,,) - mm],
where a is a fault tolerance criterion, C is a fault tolerance criterion function, 0 f is a comparison function, y is a specific value of the data record, * 328 is a predicted value and ] is a target value.
Other functional relationships can be used to specify C(a) but the percentile function is an intuitive guide to understanding why the model includes or excludes certain data records, such as Equation 2 below:
Equation (1,2) A: (((-y-324))2 = 09)
where 3 is the percentage function, 4 is the index of a record entry, and ] is
The number of log entries.
180476
-61-
Since the DOBR procedure is iterative, in one model we also specify a convergence criterion which is set in this discussion at 0.5%.
In one embodiment, given a dataset (-(404, a solution model 408 M, and an error acceptance criterion 424 OC, DOBR can be implemented to reduce bias in training the model 408 M. In
<p dir="rtl">5 In some embodiments, the solution model 408 M is implemented by an explanatory engine, including, for example, a processing device, memory, and/or storage device. According to one embodiment, the explanatory methodology computes model coefficients, ()0 402 and model estimates 13 410 for all records that apply the solution model, 408 M, to the entire input data set [-(+ 404 according to, for example, Equation 3 shown below:</p>
<p dir="rtl">10 Equation C</p>
Where 0 indicates an initial state, and no indicates an input record.
Then, according to an illustrative model, the total error function 418 calculates the total error of the initial model 20 according to, for example, equation 4 below:
Equivalent 4 - 55'
<p dir="rtl">15 where 55 is the total error of the initial model and 0 indicates the initial value.</p>
Then, according to an illustrative model, the error function 412 calculates the model errors according to, for example, Equation 5 below:
Equation 5 E 2, — 7-), E (:L)) - let},
where E is the predicted registration error, and * denotes the frequency of register selection.
<p dir="rtl">20 Then, according to an illustrative model, the error function 412 computes a new data record selection vector for * according to, for example, Equation 6 below:</p>
180476
-62-
Equation 6
where a is the log selection vector.
Next, according to an illustrative embodiment, the data record selection module 414 calculates the non-anomalous data records to be included in the model calculation by selecting only the records in which the vector
<p dir="rtl">5 The registration selection is equal to 1, according to, for example, equation 7 below:</p>
Equation 7 ((1,1) e = br(,2,)
Where 1 in a given index, indicates the set of records included in DOBR as non-outlier values.
Then, according to an illustrative model, the 408 model that includes the most recent transactions calculates the 402 values.
<p dir="rtl">10 New predicted 420 and model parameters 402 from selected DOBR data records 416</p>
According to, for example, equation 8 below:
Equation: ((3,) ,**),»(m), 1(2))
Then, according to an illustrative model, Model 408 calculates, using the coefficients of the new model, the values of
New prediction 420 for the full dataset. This step reproduces the calculation of the predicted values 420
<p dir="rtl">15 For DOBR records specified in the formal steps, but in practice, the new model can be applied to DOBR records that have been removed according to, for example, Equation 9 below:</p>
1 for equation 9
Then, according to an illustrative model, the total error function 418 calculates the total error of the model according to, for example, equation 10 below:
<p dir="rtl">20 Equation 10 (-0,,92,41) = Lb<sup>G</sup>،</p>
Where 9 is the target output.
180476
-63-
Next, according to an illustrative model, the convergence test 424 tests the convergence of the model according to, for example, Equation 11 below:
where β is a convergence criterion 422, e.g., 0.5%.
<p dir="rtl">5 In some embodiments, the convergence test 424 may terminate the iterative process if, for example, the error rate is less than, for example, 0.5%. Otherwise, the process may return to the initial data set 404. Each of the above steps may then be performed and the convergence criterion 422 may be retested. The process is repeated until the convergence test 424 is less than the convergence criterion 424.</p>
<p dir="rtl">10 Figure 7 is a graph illustrating an example of the relationship between model error and error acceptance criterion for another computer-based explanatory machine learning model with low bias according to one or more of the current disclosure models.</p>
Because it is not an input parameter to DOBR and model results can vary based on the value selected, in a model, it is important to document a data-driven procedure to justify the value.
<p dir="rtl">15 used. In practical applications where DOBR has been developed and applied, there is no theoretical basis (yet) for its selection. However, in practice, a plot of model error versus OC can produce a change in slope where the apparent effects of outliers are reduced. Figure 1 illustrates such a plot for calculating a nonlinear regression 402 related to the power generation criterion according to the embodiment of the present invention.</p>
In one model, the general shape of this curve is predetermined in that it will always start with the largest error.
<p dir="rtl">20 At OC = 100%, the model error is zero at OC = 0. In Figure 7, notice that the slope of the slope changes when OC is about 85%. For all smaller values of OC, the slope is almost constant. The change in slope at this point indicates that the variance of the model does not change with the removal of data records, or in other words, there are no outliers at these levels to accept error. When OC is</p>
180476
-64-
Greater than 85%, there are at least two apparent changes in slope indicating that some parts of the data set have behaviors or phenomena that are not accounted for in the model. This visual test can help set the appropriate error tolerance level and also determine whether DOBR is needed at all. If the slope of the line in Figure 7 does not change, the model is
<p dir="rtl">5 Satisfactory variability in the data. No outliers in the models and no DOBR required.</p>
In simulation studies where certain amounts of additional variance have been added to a data set, curves similar to those of Figure 6 show an initially steep slope line that intersects the slope of a rented value at approximately the error tolerance value programmed into the simulation. However, in practice, when outliers are observed, the transition to a flat slope generally occurs gradually, indicating that there is more
<p dir="rtl">10 One type of variance that is not accounted for in the model.</p>
Calculating an appropriate error tolerance value is an essential part of using DOBR as it visually shows the magnitude and severity of the anomalous effects on the model results. This step documents the choice of ∝ and can justify not using DOBR if the anomalous effect is judged to be small compared to the value of the model predictions from the anomalous data.
<p dir="rtl">15 In some models, ∝ and the model error versus ∝ can be used as a measure to determine the best performing model or consistent set of models for a given scenario. Given the data sets</p>
Different models may vary in degree of abnormality, the exact value of ∝ for the data and for the model may change the performance of the model. Hence, model error can be used as a function of the error tolerance level to determine the degree to which a particular model can account for the variance in the data by having model error indicate a greater or lesser amount.
<p dir="rtl">20 Tolerance to data variance in order to make accurate predictions. For example, the accuracy of model predictions can be adjusted by choosing a model and/or model parameters that exhibit, for example, a low model error tolerance value to a high error tolerance value to choosing a model that is more tolerant to data anomalies.</p>
In some embodiments, model selection may be automated using, for example, rule-based programming and/or machine learning models to determine the best performing model for a data set according to the equilibrium
18476
-65-
Model error and error tolerance criteria. Then, a model can be automatically selected that best represents the outliers in the dataset. For example, the model error can be compared across models for one or more error tolerance values, with the model with the lowest model error automatically selected for generating predictions.
<p dir="rtl">5 As a result, DOBR machine learning techniques in accordance with aspects of the present disclosure provide more efficient model training.</p>
Effectiveness, as well as improved visibility into data and model behaviors for an individual dataset. As a result, in areas such as artificial intelligence, data analytics, business intelligence, and others, machine learning models can be more effectively and efficiently tested on different types of data. Model performance can then be more efficiently evaluated to determine the optimal model for the application and type
<p dir="rtl">10 Data. For example, AI applications can be optimized with models selected and trained using DOBR for the type of intelligence being produced. Similarly, business intelligence and data analytics, as well as other applications such as physical behavior prediction, content recommendation, resource usage predictions, natural language processing, and other machine learning applications, can be optimized using DOBR to tune model parameters and select models based on anomaly characteristics and model error in response.</p>
<p dir="rtl">15 To outliers.</p>
Figure 8 is a graph illustrating an example of the relationship between model error and error acceptance criterion for another exploratory computer-based machine learning model with low bias according to one or more of the exemplars of the present disclosure.
According to an example of DOBR model on a dataset, we use a compressive strength dataset.
<p dir="rtl">20 Concrete 504 downloaded from the California-Irvine Machine Learning Data Repository. This dataset contains 1030 observations or records or states with 8 independent variables. The first seven describe the composition of the concrete at the age given in days: amount of cement, superplasticizer, blast furnace slag, coarse aggregate, fly ash, fine aggregate, water, and age.</p>
18476
-66-
The resulting variable is the compressive strength of concrete measured in megapascals (MPa). For comparison, 1 MPa is approximately 145 pounds per square inch. A linear regression model is created according to, for example, Equation 12 below:
Equation 12 Compressive strength of concrete = +40 [=5
<p dir="rtl">5 Where: is a coefficient calculated by a linear regression model, 2 is observations for 8 variables, and t is an index variable.</p>
Figure 8 is generated by running a linear regression model 504 as a function of DOBR, , from 100 to 60%. From OC = 100% to about LM = 95% there is a rapid decrease in model error, as shown by the regression 506, and then the error reduction decreases as a function of LM = 10 at a slightly lower rate until OC = 85%. From this point onwards, OC decreases at a constant rate, as
The point at which the error begins to decrease at a constant rate is the point at which the effect of the outlier is removed from the model calculation. In this case, the selection point is >= 85%.
In one model, DOBR is then a linear regression model that is rerun for OC = 15 92.5% to determine the best model that fits the non-outlier data. Figure 9 and Figure 10 show the results.
These calculations are made using the full dataset 512 (Figure 9) and the DOBR version (Figure 10) that includes the identified outliers and removed from the calculation. The 516 outliers, marked with red crosses, are calculated from the non-outlier model. Both graphs show the actual target values versus the predicted target values with the diagonal lines 510 and 514, respectively Figure 9 and Figure 10, 20 which illustrate the equivalence. The calculation of the full dataset (Figure 9) illustrates how outliers can
The results are biased. The DOBR-adjusted graph (Figure 10) shows the bias removed by the diagonal line 514 that divides the non-outlier values 518 as well as clear clusters of outlier data points 516 that may require further study.
180476
Figure 9 is a graph illustrating an example of the relationship between compressive strength and predicted compressive strength for a basic computer-based machine learning model with low bias according to one or more of the embodiments of the present disclosure.
Figure 10 is a graph showing an example of the relationship between compressive strength and predicted compressive strength5 for another computer-based explanatory machine learning model with low bias according to one or more of the
Current disclosure forms.
Identifying outliers and the patterns they form in the above-mentioned type of graphs is useful for additional benefits of the DOBR method in industrial applications. Outliers can form patterns or clusters that are simply not noticed by other methods. This information is simply generated by
<p dir="rtl">10 The DOBR method uses a model provided by the analyst. No additional information or assumptions are required. In practice, the anomaly set identified by DOBR can provide useful information to refine, provide insights, or validate the underlying model.</p>
Figure 1 1 is a block diagram of another illustrative computer-based system for machine learning predictions using DOBR in accordance with one or more embodiments of the present disclosure.
<p dir="rtl">15 In one embodiment of the present invention, a machine learning procedure begins with a dataset, * consisting of independent variables η, m records of length, and an m-value matrix X) of target variables, . In one embodiment, to train a machine learning model, the dataset (5,) is split into two randomly selected subsets of a predetermined size: one to train the model and one to test its predictive accuracy, e.g., Equation 13 below:</p>
<p dir="rtl">20 Al-Ma'asala 13 is deception}- {Rum</p>
where * is a subset of the independent variables X of the data set, and n is a subset of the independent variables * of the data set.
180476
-68-
For this discussion, a 70%/30% split is used for training (η records) and testing (g records) (i.e., 70% of the records are trained and 30% are tested), however any appropriate split can be used, e.g., 50%/50%, 60%/40%, 80%/20%, 90%/10%, 95%/5%, or another appropriate training/test split.
<p dir="rtl">5 A machine learning model, 1, is tested using (25) which is tested by computing a set of predicted target variables 32) expressed as shown in, for example, Equation 14 below:</p>
Equation 14 x8lm«, ίτπ (wd '£)] x 41,,2).
In an illustrative model, the accuracy of the model is measured as a criterion, 9 which might, for example 10 , have the following form:
The equation 15 (9-j) 224 = 1 8;d 8', d 1.
In an illustrative model, in training and testing environments, we can measure outliers directly since we have both input and output variables. In general, outliers in model predictions, such as large deviations from the actual target variable values, are due to the model's inability to transform
<p dir="rtl">15 Input values are set to predictive values close to the known target variable. The input data for these records contain the effects of factors and/or phenomena that the model cannot map to reality as defined by the target variables. Keeping these records in the dataset can bias the results since the model coefficients are calculated assuming that all data records are equally valid.</p>
<p dir="rtl">20 In some embodiments, the DOBR process described above, for example, in reference to Figure 6 above, works for a given data set where the analyst wants the best model to fit the data by removing outliers that bias the results in an adverse way. This increases the predictive accuracy of the model by restricting the model solution to a subset of the initial data set.</p>
180476
-69-
from which outliers have been removed. In an illustrative example, the DOBR auxiliary solution has two output results:
a) A set of X values, model parameters, and model solutions for which the model describes the data, and
<p dir="rtl">b) A set of X values, model parameters, and model solutions for which the model does not describe the data.</p>
<p dir="rtl">5 Therefore, in addition to computing a more accurate model for the constrained dataset, in the models, DOBR also provides an outlier dataset that can be further studied for the given model to understand the reason or reasons for the model's high prediction error.</p>
In an illustrative example of a machine learning architecture as described earlier in this section, the predictive model is computed from the training data and this model alone is used in the testing phase. Since,
<p dir="rtl">10 During the design, the testing phase may not use target values to identify outliers, and the DOBR methodology described above in reference to Figure 6 may not apply. However, there is an illustrative aspect of the DOBR methodology that may not have been used above: the possibility of an outlier-non-outlier classification as suggested by the DOBR output results mentioned earlier.</p>
To describe DOBR in a machine learning application of one of the embodiments of the present invention, the dataset may be divided
<p dir="rtl">15 Into two randomly selected parts: one for training and one for testing. In the training phase, both the independent and target variables are kept, but in the testing phase the target variables are hidden and the independent variables are used to predict the target variable. Only the known target variable values are used to measure the prediction error of the model.</p>
In one embodiment, given a training dataset of 1 1,3 604 with Π records,
<p dir="rtl">20 A machine learning model 608 L, and an error tolerance criterion 622 OC, DOBR may be implemented to reduce bias in training the machine learning model 608 L. In some embodiments, the machine learning model 608 L is implemented by a model engine, including, for example, a processing device, memory, and/or storage device. According to one embodiment, the explanatory methodology model estimates 0:1 606 for all records that apply</p>
180476
-70-
Machine learning model 1 608, to the full input dataset *»l 1,9 604 according to, for example, equation 16 below:
32 )and — ί^Χ,Ύύίτώι,Χίταήι] 16 Equation
Where 0 indicates an initial state, and X indicates an input record.
5 Then, according to an illustrative model, the total error function 618 calculates the total error of the initial model 20 according to, for example, equation 17 below:
The equation 17 || No,'m'j;,<sup>0</sup>II -55;
where 20 is the total error of the initial model.
Then, according to an illustrative model, the error function 612 computes the model errors according to, for example,
10 Example, for equation 18 below:
Equation 18(12)<sub>k</sub> - ), E)) - (*]،
where E is the predicted log error, and k denotes the frequency.
Then, according to an illustrative model, the error function 612 computes the selection vector of the new data record according to, for example, Equation 19 below:
15 00.00 'no'';w't.
where I is the log selection vector.
Next, according to an illustrative embodiment, the data record selection unit 614 calculates non-anomalous data records to be included in the model calculation by selecting only records for which the record selection vector is equal to 1, according to, for example, Equation 20 below:
20)
Equation 20
,2) = ((2,90),.:,1 E (1,1
180476
-٩١-
where in in index refers to a set of DOBR records included as non-outlier values.
Next, according to an illustrative embodiment, the terminal machine learning unit 608 with the latest coefficients 602 calculates the new predicted values 620 for the complete training set 604 using the specified DOBR data records according to, for example, Equation 21 below:
5 Equation 21 = ,( ]03
Then, according to an illustrative model, the total error function 618 calculates the total model error according to, for example, equation 22 below:
The fair 22 ||l - ة*•،
Next, according to an illustrative model, the 624 convergence test tests the convergence of the model according to, for example,
<p dir="rtl">10 Example, for equation 23 below:</p>
The equation c0>b,
where 0 is a criterion of convergence of 622, for example 0.5%.
In some embodiments, the convergence test 624 may terminate the iterative process if, for example, the percentage error is less than, for example, 0.5%. Otherwise, the process may return to 15 training data sets 604.
In some models, the DOBR iteration measures how well the model predicts itself rather than its accuracy on a test data set. The goal here is to test the model's ability to predict the target variable and records with large deviations are systematically removed to improve the model's ability to focus on the vast majority of data where the data predictions are relatively good.
<p dir="rtl">20 This process is done on the same data set. It does not make sense to remove records from the training set if outliers are identified in the test set. This process is fundamental to the DOBR method as</p>
180476
-٩٦-
Records that were removed in a previous iteration are reinserted after a new model (new model parameters) is calculated. This process requires the use of the same dataset.
In one embodiment, this iteration is performed after the learning model has been defined. Based on the problem to be solved, in one embodiment, the user selects a machine learning algorithm and then specifies specific hyperparameters.
<p dir="rtl">5 which “tune” or initialize the model. These parameters can be chosen using standard techniques such as cross-validation or simply by plotting the test error as a function of user-supplied parameter ranges. The specific values used can improve prediction accuracy versus computation time while ensuring that the model is neither over-fitted nor under-fitted. There are many powerful tools to assist in this process but user experience and intuition are also</p>
<p dir="rtl">10 Valuable advantages in choosing the best hyperparameters for the model. Below are discussed the selected models and their associated hyperparameters used in the examples.</p>
The error tolerance versus model error plot is calculated from this step by applying a sequence of error tolerance values and tabulating or plotting the results. These plots identify the portion of the data set that is an outlier in the sense that the error contribution to it is marginally greater than the error contribution to
<p dir="rtl">15 Data records that fit the model. In practice, too, these plots can show more than one type of variance that is not explained by the model. The slope can vary as it converges to the model slope. These variations can help in investigating the nature of additional behavior encoded in the data that is not explained by the model. Records that occupy different slope intervals can be identified and their further examination can provide insights that may help in constructing a more robust model.</p>
20 power.
In one model, when training, as shown above, two models were computed:
Model 1
where (/03 is a reference model used as a basis for measuring accuracy improvements; and
180476
-13 -
Model 2
where 021 is the base DOBR model, built from the closely clustered censored anomaly records and trained on non-anomaly data (3,+).
<p dir="rtl">5 In the models, the errors associated with Model 1 and Model 2 are, for example, 1:(03-9 a*2:,0221 1 22 respectively.</p>
Hence, in the models, the basic model 21 suggests that it may be a better predictor of non-abnormal records. However, the test data set is uncensored, containing both non-abnormal and abnormal values. Therefore, it is uncertain whether the application of the model
<p dir="rtl">10 A non-outlier specific to uncensored test data will produce a better predictive model than a 22/3). However, in many cases, 22 can be observed in many cases to be statistically equal or greater than</p>
In non-machine learning applications where the goal is to compute the best predictive model for a given dataset, a DOBR model, computed from selected (non-outlier) records, always produces a model error less than 15 since the selected outliers are removed. In the constrained case of no outliers, the DOBR model error is equal to the total model error since the datasets are the same.
However, in machine learning applications, the goal may be to develop a model using a subset of the available data (training) and then measure its predictive accuracy on another subset (testing). However, in some models, the DOBR methodology removes outliers from the model at each iteration before computing
<p dir="rtl">20 Model parameters. When developing a machine learning model, this can be done in the training phase, but by definition, the target values in the test can only be used to measure the predictive accuracy of the model without advanced knowledge of outliers. This observation means that the standard DOBR methodology can be generalized by using more DOBR model information computed in the training phase.</p>
180476
-74-
Figure 1 1 is a block diagram of another illustrative computer-based system for low-bias machine learning according to one or more embodiments of the present disclosure.
In models, when training, as described above, the following information is produced: Training data set selected by DOBR for non-outlier values (2,73) Training data selection vector
<p dir="rtl">5 DOBR for non-outliers and DOBR selected training dataset values for outliers 2(35(2) and DOBR selected training dataset vector for outliers dir l}.</p>
In the models, DOBR classifies the training data into two mutually independent subsets. In addition, we also have the corresponding selection vectors that provide a binary value: a classification value (non-outlier or outlier) for each record in the training dataset, for example 10, according to Equation 24 below:
Equation 24
<p dir="rtl">-- + „;3)- •11 0.: = ί,) + t4(,2) = :-(2,3):<sub>5af</sub>,w/iere(x,y)<sub>B</sub>&(x,y)٠</p>
In the models, the full set of training data features, ,2 and the generated DOBR labels, ,0 are used to build/train a machine learning model for the classifier, e.g., the stored
15 In the model library 303. This model is applied to the test dataset, 2t20%, to classify the test data records as outliers or non-outliers based on the knowledge established by the DOBR training database. For example, the machine learning model of the classifier is implemented according to Equation 25 below:
Equation 25 [-,(,1.,,,,.,])] - 1]].
<p dir="rtl">20 Hence, in one model, 01] produces two predictive test datasets; , where 0 or , respectively. The above information generates many possible predictive models for the full dataset" for analyzing the test dataset. In some models, the three that showed the most predictive improvements for the full dataset are:</p>
180476
-15 -
Model 3
Form 4
iv,
,,0(2,)] = mm-s-l
H2Hhdh,l-„،^^0-l-„،-„،^2'
Form 5
{,la-l-ay-a-l_„,^la-ay = ag
In some embodiments, for (1, the machine learning model 1 608 is defined by the non-anomalous data, (3,+) and is applied to a test data labeled with in DOBR to predict the non-anomalous test values. The same procedure is performed for the anomalous data. In the models, the goal of this combination is to 10 use the most accurate predictive model with its corresponding data set. In other words, this tests
The overall predictive accuracy of non-abnormal and abnormal models applied separately to their respective datasets defined by the DOBR classification.
In some embodiments, for (2), the machine learning model 1 608 is defined by the data
Training, and also applied to the DOBR-labeled X test data. 15 This model uses the extensive knowledge T ««*»,((*)t to predict target values for outlier and non-outlier X values.
DOBR knowledge. The purpose of this model is to test the predictive accuracy of the fully trained model applied separately to non-anomalous and anomalous datasets labeled with DOBR.
In some embodiments, the third model (2) is a hybrid that combines the predictive properties of the previous two approaches.
This model tests the predictive usefulness, if any, of linking the 2,(37(2)] model 608 that was trained.
<p dir="rtl">20 On the overall training B2(77,)7 the selected model trained on DOBR-classified outliers in the training set applied to their respective labeled datasets. There are additional hybrid models that can be explored with further research.</p>
180476
-6 ٦-
In each of these three models and others, the full test dataset is predicted using both anomalous and non-anomalous records labeled with DOBR. The ability of DOBR to improve the overall predictive accuracy of a machine learning model is tested using these models. However, the primary benefit of DOBR is to identify and remove outliers from the model and compute the best predictor for the model.
<p dir="rtl">5 From the remaining non-outliers. By definition, outliers identified by DOBR are records that contain variation that is not adequately described by the current variables (or features) that gave rise to the machine learning model used.</p>
In some embodiments, using imputed anomalous or non-anomalous data sets, the analyst has at least three or more options. In one embodiment, the first option is to apply the model
<p dir="rtl">10 Basic, (/e), and do not apply DOBR. A data-driven strategy is used when the risk-to-error model acceptance curve is close to a linear relationship. In one embodiment, the second option is to apply one or more of the models: 11, (2, or 51, and sum, for example, the average of the results. In one embodiment, the third option is to develop predictions for non-abnormal records only and further investigate the abnormal data to develop a modeling strategy for the data set</p>
<p dir="rtl">15 These new specializations—for example, changing a machine learning model or adding variables to account for unexplained variance, etc.</p>
Regarding option 3, there are several ways to calculate the non-outlier dataset and two possible options are mentioned here. One reason for the relatively large number of possibilities may be due to the non-linearity of many of the machine learning models applied. In general,
20 [0 to 1,0.(7,) not [2,0,(3,32)11!A. This discrepancy may be due to the complexity of
Many machine learning models. The equality applies to linear regression, for example, but not as a general rule for machine learning models.
In models, for non-abnormal predictions, the DOBR method was not initially designed to optimize the prediction of the full data set. By design, the method converges to the best set of values.
<p dir="rtl">25 Anomalies based on the model and available data set. The remaining data and model calculations provide accuracy.</p>
180476
-٦٦-
Improved but no guidance on how to make predictions for outliers. The implicit decision is to apply a different model to the outlier dataset that reflects unique differences in the data not present in the non-outlier model.
In the models, two models are specified to test the accuracy of the non-outlier prediction—removing outliers from the analysis.
<p dir="rtl">5 The first option for selecting a non-outlier dataset applies a classification vector, DOBR1, to the reference model, (/3) and, for example, to model 6 below:</p>
Form 6
() 01 = [-,,,,(.3,) 01 = (,t
In models, the reference model uses the full model specified by the training data to make predictions from
<p dir="rtl">10 Dataset, 8•2. Next, a classification vector is applied to remove the predicted outliers based on the DOBR knowledge obtained from the training dataset. This model applies DOBR to the more general domain model or the broad domain model.</p>
In the models, the second model applies DOBR in the most "accurate or accurate" way using the DOBR model generated from the training phase from the non-anomalous training data, on the records.
<p dir="rtl">15 Selected by the classification model only, not, for example, according to Model 7 below:</p>
Form 7
There are other models that can be formed from the analytical formulas developed in this research, and depending on the problem, they may have great potential to improve the predictive ability. However, the models
<p dir="rtl">20 Used here, OtJ and {Os are limiting cases that represent broader and narrower versions in terms of using the training domain and specifying the model.</p>
180476
-78-
In the models, to test the predictive accuracy of the developed DOBR models specified above, e.g., Models 3-7, we use /3) as a comparison baseline for the models (*, (2, (5 (Models 3, 4, and 5, respectively). For 4 and 7 (Models 6 and 7, respectively), and the model predictions for the non-outlier dataset, the comparison baseline is 17]. Hence, in the models, it can be determined
<p dir="rtl">5 The error according to, for example, equations 26, 27, and 28 below:</p>
Equation 27
<p>= To<sub>0</sub>Let the exponent II = 2(22)' where m - length of dataset</p>
Equation 280
1
<sub>cheat</sub>Oh, yeah.<sub>A</sub>D<sub>1</sub>=2=))=24(-/,)22 u
Equation 29
• The genera11 = 2
Dataset is confusing - ) =),. A .. -
:0-- A '
In the following illustrative model examples, the DOBR predictive accuracy is measured by (if applicable) 21, 52, and/or ≥2. For non-outlier dataset errors, 4 and ≥4, the measure of improvement is the error reduction relative to the adjusted outlier baseline error. The modification described below is described in relation to the example results.
<p dir="rtl">20 In some examples of DOBR illustrative machine learning improvements, the accuracy of the five predefined models can be tested using seven machine learning regression models: linear regression, support vector regression, Lasso, neighborhood regression, bagging regression, and random forest regression. The machine learning regression models are examples of a wide range of model architectures. Additional models or</p>
180476
-79-
Alternatives, such as neutral networks, clustering, and consistent set models, among others and combinations thereof.
Linear regression is a method that gives analysts insights into a process where the coefficients (or model parameters) can have technical/process-related meaning. One of the process models, 5 represented by an equation, must be provided by the analyst and the coefficients are determined by minimizing the error between the target values.
Predicted and data-driven.
LASSO, short for “least absolute shrinkage and selection factor,” is a regression-related methodology where an additional term is added to the objective function. This term is the sum of the absolute values of the regression coefficients and is minimized according to a parameter provided by the user. The purpose of this additional term is to add a penalty for increasing the value of the variable coefficients (or feature). It does not retain
Minimizing the prevailing coefficients can help reduce the effects of covariance or collinearity (or feature) that are difficult to interpret.
Decision tree regression can mimic human thinking and is intuitive and easy to interpret. The model chooses a decision tree structure that logically explains how the values of x produce the target variable. The
<p dir="rtl">15 Parameters such as maximum depth and minimum samples per sheet are specified by the analyst in the testing/training machine learning exercise.</p>
Random forest regression is based on the decision tree method. Just like forests are made up of trees, a random forest regression model is made up of groups of decision trees. The analyst specifies the forest structure by supplying the units of estimation (number of trees in the forest), some parameters similar to the decision trees with maximum tree depth, leaf characteristics, and technical parameters related to the
How to calculate and apply model error.
k-NN refers to k-nearest neighbor methods where the predicted value is calculated from the k-nearest neighbors in the range (or feature) x. The metric for measuring the distance and the specified number of nearest neighbors are
18476
-80-
To use are the key parameters that the analyst has specified in tuning a model for predictions on a given data set. This is a straightforward method that can work well for regression and classification predictions.
Support vector regression is a versatile machine learning method that has many variations. Regression refers to the fit of a model to the data and the improvement is usually the minimum error between
<p dir="rtl">5 The predicted variable and the target variable. Using support vector regression, the error criterion is generalized to assume that if the error is less than a certain value "p", we say that this is a good enough approximation and only errors greater than "p" are measured and improved. In addition to this feature, the method allows the data to be transformed into non-linear ranges by the criterion or, in some cases, user-defined transformation functions or kernels. The multidimensional data structure is used where the goal is to compute predictions</p>
<p dir="rtl">10 Powerful - not for modeling data in the spirit of traditional regression.</p>
Bagging regression computes prediction estimates from drawing random subsets with replacement. Each random sample computes a decision tree prediction (by default) for the target variable. The final consistent ensemble prediction value can be computed in several ways—the mean value is one example. The basic machine learning variables are the number of estimators in each consistent ensemble, the number of variables (or features),
<p dir="rtl">15 and samples to be drawn to train each estimator, and selection/replacement guidelines. The method can reduce variance compared to other methods such as decision tree regression.</p>
The classifier model, //,12:2 is an illustrative example because it is applied to DOBR non-outlier/outlier classifications and training set values X to identify non-outlier and outlier values in the test dataset. This is a crucial step in the machine learning application of DOBR because it conveys
<p dir="rtl">20 Identify outliers from the training set to the test or production dataset. If there are inappropriate classifications, the benefit of the DOBR methodology to improve the accuracy of machine learning predictions will not be realized.</p>
Decision tree, k-NN, random forest, and bagging classifier models were tested for their classification accuracy. The bagging and random forest models were selected and both models were tuned to produce the correct error tolerance fraction for non-outlier values. A more detailed classification analysis can suggest other models.
180476
-81-
A comprehensive analysis of the classifier, although the accuracy of the classification is important, is beyond the scope of this preliminary discussion.
Figure 12 is a graph showing an example of the relationship between model error and error tolerance for some computer-based explanatory machine learning models that have low bias for strong prediction.
<p dir="rtl">5 Concrete according to one or more embodiments of the present disclosure.</p>
The first example uses the same dataset as described above with reference to concrete compressive strength, where DOBR is applied to the entire dataset. As a short summary, this dataset contains concrete compressive strength as a function of its composition and exposure as defined by 8 quantitative input variables. The dataset contains 1030 records or instances and can be found 10 in the UC Irvine Machine Learning Repository archive.
Machine learning training splits this dataset into 70%:30% with model tuning performed on the training dataset (70%) and prediction results measured using the testing dataset (30%).
The model tuning results of seven machine learning models in predicting concrete compressive strength are shown in Table 15.1 below.
Table 1
<tr><td><p>fit_intercept=False, normalize=False</p></td><td><p dir="rtl">linear regression variable</p></td></tr><tr><td><p>alpha=4, fit_intercept=False</p></td><td><p>LASSO</p></td></tr><tr><td><p>max_depth=6, min_samples_split=2</p></td><td><p dir="rtl">Decision tree regression variable</p></td></tr><tr><td><p>n_estimators=3, min_samples_leaf=30</p></td><td><p dir="rtl">Random Forest Regression Variable</p></td></tr><tr><td><p>n_neighbors=3</p></td><td><p dir="rtl">K-neighborhood regression variable</p></td></tr>
18476
-82-
<tr><td><p>C=10, gamma=0.0005, kernel='rbf'</p></td><td><p>SVR</p></td></tr><tr><td><p>n_estimators=25, max_samples=35</p></td><td><p dir="rtl">regression variable for filling</p></td></tr>
The default model parameters are not specified (for example, for Python 3.6) because they do not add information to the results. In models, the tuning process is an exercise in choosing parameters that minimize the errors of the training and test datasets using the mean squared error as an indicator. More complex algorithms can be applied but the straightforward approach was used simply to ensure that the results are not overfitted or underfitted to the dataset error.
In one embodiment, for the DOBR application, the percentage of data, if any, where the error is excessively large is determined. In the models, machine learning models are applied to a series of error tolerances that record the corresponding model errors. This is done only for the training dataset as the test dataset is used only to measure the accuracy of the machine learning model's prediction. The percentage of data 10 listed in the model, "error tolerance", indicates the amount of total error in the model that the user desires
In its acceptance it also refers to the part of the data that the model adequately describes.
In the models, the percentage of error acceptance sequence ranges from 100% to 60% in increments of 2.
Figure 13 is a graph illustrating an example of the relationship between model error and error tolerance criterion for some low bias computer-based explanatory machine learning models for prediction using 15 energy according to one or more of the models of the present disclosure.
The second example contains appliance energy usage data along with environmental and home lighting conditions sampled every 10 minutes for 4.5 months. It includes 29 attributes: 28 input variables, one output (target variable), and 19,735 records: The dataset and documentation can be found in the UC Irvine Machine Learning Repository archive.
<p dir="rtl">20 Similarly to the above, in the models, the model tuning results of seven machine learning models in device power usage prediction are shown in Table 2 below.</p>
18476
-83-
Table 2
<tr><td><p>fit_intercept=False, normalize=False</p></td><td><p dir="rtl">Linear regression</p></td></tr><tr><td><p>alpha=4, fit_intercept=False, max_iter=100000, tol=0.01</p></td><td><p>LASSO</p></td></tr><tr><td><p>max_depth=22, min_samples_leaf=2</p></td><td><p dir="rtl">Decision tree regression variable</p></td></tr><tr><td><p>n_estimators=6</p></td><td><p dir="rtl">Random Forest Regression Variable</p></td></tr><tr><td><p>n_neighbors=9</p></td><td><p dir="rtl">K-neighborhood regression variable</p></td></tr><tr><td><p>C=1000, gamma=0.001, kernel='rbf'</p></td><td><p>SVR</p></td></tr><tr><td><p>n_estimators=20, max_samples=15</p></td><td><p dir="rtl">regression variable for filling</p></td></tr>
In the models, the parameters of the hypothetical model are not specified (e.g., for Python 3.6) because they do not add information to the results. The tuning process was an exercise in choosing parameters that minimize the errors of the training and test dataset using the mean squared error as an indicator.
<p dir="rtl">5 More complex algorithms but a straightforward approach is simply used to ensure that the results are not overfitted or underfitted to the data set error.</p>
In one embodiment, for the DOBR application, the percentage of data, if any, where the error is excessively large is determined. In the models, the machine learning models applied to the sequence of fractional acceptance errors record the corresponding model errors. This is done only for the training dataset where
<p dir="rtl">10 Use only the test dataset to measure the prediction accuracy of a machine learning model. The percentage of data included in the model, “error acceptance,” indicates how much of the total error in the model the user is willing to accept and also indicates the portion of the data that the model adequately describes.</p>
In the models, the error acceptance ratio ranges from 100% to 60% in increments of 2.
18476
-84-
Figure 12 and Figure 13 illustrate, in part, the ability of machine learning models to adapt to highly variable data. The closer the lines are to being (straight) lines, the better the model is able to adequately describe the data, which translates into fewer outliers, if any. The linear behavior of many models applied to concrete data shows that they can describe the entire training dataset.
<p dir="rtl">5 Almost adequately. The nonlinearity of the energy dataset results indicates that there is a large proportion of data records where the models produce inaccurate predictions or outliers.</p>
For each curve in the concrete data plot above, including, for example, linear regression 530, LASSO 540, decision tree regression 522, random forest regression 528, k-neighborhood regression 524, support vector regression (SVR) 520, and packing regression 526, and in the 10 energy use data plot above, including, for example, linear regression 730,
740 LASSO, decision tree regression722, random forest regression728, k-neighbor regression724, support vector regression720 (SVR), and bagging regression726, the straight line defined by the low error tolerance ratios can be interpolated to determine the error tolerance value where the outlier fraction begins, according to the current modeling limit. This process can be automated but in practice, it can be done15 manually to ensure that the chosen error tolerance value reflects the analyst's judgment.
The process of extrapolating and choosing a percentage of error tolerance is relatively simple but has very important implications. It indicates how well the proposed model fits the data. The complement of the error tolerance value is the percentage of the dataset that is anomalous, that is, the percentage of records in which the model fails to make relatively accurate predictions. This is important information in choosing a machine learning (or any) model for a given dataset and practical application. Table 3 represents the error tolerance values chosen for each mode for two sets of
Data.
Table 3
<tr><td><p dir="rtl">Device power usage</p></td><td><p dir="rtl">concrete pressure</p></td><td></td></tr>
18476
-85-
<tr><td><p>%84</p></td><td><p>%80</p></td><td><p dir="rtl">Linear regression</p></td></tr><tr><td><p>%84</p></td><td><p>%80</p></td><td><p>LASSO</p></td></tr><tr><td><p>%90</p></td><td><p>%94</p></td><td><p dir="rtl">tree of decision</p></td></tr><tr><td><p>%90</p></td><td><p>%90</p></td><td><p dir="rtl">Random Forest</p></td></tr><tr><td><p>%84</p></td><td><p>%88</p></td><td><p dir="rtl">The neighborhood is closer to k</p></td></tr><tr><td><p>%84</p></td><td><p>%94</p></td><td><p dir="rtl">support vector</p></td></tr><tr><td><p>%84</p></td><td><p>%92</p></td><td><p dir="rtl">Packing</p></td></tr>
In the models, the predictive accuracy of the selected DOBR values is only compared to the reference model. This is a fundamental tool of DOBR since the method itself does not provide any specific information about the increase in accuracy.
Predictive for the entire dataset. Therefore, DOBR analysis presents the analyst with a potential trade-off: to have better predictive power for a portion of the dataset but not provide information for outliers.
<p dir="rtl">5 The question addressed in this section is how accurate, if at all, are the specified DOBR results compared to the predictions of the corresponding reference model test data.</p>
The reference error is calculated for the full data set. Adjusted reference error values for comparison to non-outlier data sets are calculated by multiplying the full reference error by the error tolerance value. For example, if the reference error is 10.0 and the error tolerance value is 80%, then the adjusted reference error is 10 × 80% or 8.0. The interpretation uses the definition of "error tolerance". If
If the data is non-abnormal on 80% of the data, for example, 80% of the total error should remain in the non-abnormal data. This is the definition of error tolerance.
The results measuring the performance of the predictive accuracy for selected non-abnormal DOBR values are shown in Table 4 and Table 5 below, corresponding, for example, to the concrete strength dataset and the 15 energy dataset, respectively. The reference error is calculated by multiplying the actual error acceptance percentages by
18476
-86-
Point estimates: 7721.3. The random forest classifier was not applied to the concrete dataset as it was judged that it would not change the conclusions regarding the improvement of predictive accuracy for
.DOBR
For all of the following statistics, results show the mean ± 95% confidence interval of 100.
<p dir="rtl">5 Selection of subsets of training and test data. In some examples in the following tables, support vector results have been calculated from a smaller number of iterations (5 or 10) to manage computational time issues.</p>
Table 4
<tr><td><p dir="rtl">ratio</p></td><td><p dir="rtl">ratio</p></td><td><p><sub>to</sub>}III</p></td><td><p>}, }II</p></td><td><p dir="rtl">Error</p></td><td><p>111</p></td><td><p dir="rtl">acceptance</p></td><td></td></tr><tr><td><p dir="rtl">The centenary</p></td><td><p dir="rtl">The centenary</p></td><td></td><td></td><td><p dir="rtl">Reference</p></td><td></td><td><p dir="rtl">Error</p></td><td></td></tr><tr><td><p dir="rtl">K</p></td><td><p dir="rtl">To improve</p></td><td></td><td></td><td><p dir="rtl">Y</p></td><td></td><td><p dir="rtl">Actual</p></td><td></td></tr><tr><td><p dir="rtl">To be envied</p></td><td><p dir="rtl">Z</p></td><td></td><td></td><td></td><td></td><td><p>(%)</p></td><td></td></tr><tr><td><p dir="rtl">N</p></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td><p>[)</p></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td><p>8.4</p></td><td><p>23.2</p></td><td><p>±7.80</p></td><td><p>±6.54</p></td><td><p>8.52</p></td><td><p>± 10.49</p></td><td><p>801.2</p></td><td><p dir="rtl">The decline</p></td></tr><tr><td><p>%</p></td><td><p>%</p></td><td><p>0-10</p></td><td><p>0.08</p></td><td></td><td><p>0.07</p></td><td><p>0</p></td><td><p dir="rtl">R</p></td></tr><tr><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td><p dir="rtl">linear</p></td></tr><tr><td><p>7.9</p></td><td><p>23.6</p></td><td><p>±7.90</p></td><td><p>±6.55</p></td><td><p>80.58</p></td><td><p>± 10.51</p></td><td><p>81.6</p></td><td><p>LASS</p></td></tr><tr><td><p>%</p></td><td><p>%</p></td><td><p>0.08</p></td><td><p>0.07</p></td><td></td><td><p>0).0)7</p></td><td><p>%</p></td><td><p>0</p></td></tr><tr><td><p>1.6</p></td><td><p dir="rtl">%0٠5</p></td><td><p>±7.45</p></td><td><p>±7.54</p></td><td><p>1.5</p></td><td><p>±7.89</p></td><td><p>96.0</p></td><td><p dir="rtl">Tree</p></td></tr><tr><td><p dir="rtl">٥/٠</p></td><td></td><td><p>0.11</p></td><td><p>0.10</p></td><td></td><td><p>0.10</p></td><td><p>%</p></td><td><p dir="rtl">Decisions</p></td></tr><tr><td><p>4.2</p></td><td><p>-</p></td><td><p>±7.95</p></td><td><p>±8.36</p></td><td><p>8.30</p></td><td><p>±9.04</p></td><td><p>91.8</p></td><td><p dir="rtl">the forest</p></td></tr><tr><td><p dir="rtl">٥/٠</p></td><td><p>%0.7</p></td><td><p>0.11</p></td><td><p>0.10</p></td><td></td><td><p>0.10</p></td><td><p>%</p></td><td><p dir="rtl">Random</p></td></tr><tr><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td><p dir="rtl">K</p></td></tr>
180476
-87-
<tr><td><p dir="rtl">().5</p><p dir="rtl">٥/٠</p></td><td><p>%3.2</p></td><td><p dir="rtl">٦.٦٦±</p><p dir="rtl">0.10</p></td><td><p>±8.44</p><p>0.12</p></td><td><p>8.18</p></td><td><p>±9.26</p><p>0.10</p></td><td><p dir="rtl">88.3</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Nearest Neighbor</p><p>kl</p></td></tr><tr><td><p>1.</p><p>0</p></td><td><p>01.4</p></td><td><p>±8.57</p><p>0.10</p></td><td><p>±9.11</p><p>0.11</p></td><td></td><td><p>±9.84</p><p>0.11</p></td><td><p dir="rtl">93.9</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Museum</p><p dir="rtl">Support</p></td></tr><tr><td><p dir="rtl">0.7</p><p dir="rtl">٥/٥</p></td><td><p>%7.3</p></td><td><p>±7.86</p><p>0.09</p></td><td><p>±8.37</p><p>0.10</p></td><td><p>7.80()</p></td><td><p>±9.84</p><p>0).10)</p></td><td><p>86.5</p><p>%</p></td><td><p dir="rtl">Packing</p></td></tr>
Table 5
<tr><td><p dir="rtl">ratio</p></td><td><p dir="rtl">ratio</p></td><td><p>^-1-.^.-.1-111</p></td><td><p dir="rtl">The lam {0'</p></td><td><p dir="rtl">Error</p></td><td></td><td><p dir="rtl">acceptance</p></td><td><p dir="rtl">Repetition</p></td></tr><tr><td><p dir="rtl">The centenary</p><p dir="rtl">To be envied</p><p dir="rtl">N</p></td><td><p dir="rtl">The centenary is envied</p><p dir="rtl">N</p><p dir="rtl">{and</p></td><td></td><td></td><td><p dir="rtl">The meadow</p><p dir="rtl">powerless</p></td><td></td><td><p dir="rtl">Actual error (%)</p></td><td></td></tr><tr><td><p>27.4</p><p>%</p></td><td><p dir="rtl">24.4</p><p dir="rtl">٥/٥</p></td><td><p>±59.48</p><p>0.50</p></td><td><p>±61.99</p><p>0.38</p></td><td><p>81.</p><p>96</p></td><td><p>±93.99</p><p>0.41</p></td><td><p>87.2</p><p>%</p></td><td><p dir="rtl">The decline</p><p dir="rtl">linear</p></td></tr><tr><td><p dir="rtl">27.8</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">26.3</p><p dir="rtl">٥/٥</p></td><td><p>±59.53</p><p>0.55</p></td><td><p>±60.80</p><p>0.42</p></td><td><p>82.</p><p>44</p></td><td><p>±94.87</p><p>0).37</p></td><td><p dir="rtl">86.9</p><p dir="rtl">"Yes</p></td><td><p>LAS</p><p>SO</p></td></tr><tr><td><p dir="rtl">21.1</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">12.6</p><p dir="rtl">٥/٥</p></td><td><p>±67.08</p><p>0.55</p></td><td><p>±74.25</p><p>0).52</p></td><td><p>804.</p><p>97</p></td><td><p>±92.06</p><p>0.51</p></td><td><p dir="rtl">92.3</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Decision tree for</p></td></tr><tr><td><p>11.6</p></td><td><p>14.0</p></td><td><p>±63.25</p></td><td></td><td><p>71.</p></td><td><p>±ΊΊ٠Έ</p></td><td><p>92.0</p></td><td><p dir="rtl">the forest</p></td></tr>
180476
-88-
<tr><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">٥/٥</p></td><td><p>0.49</p></td><td><p>0).40)</p></td><td><p>51</p></td><td><p>0.36</p></td><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Random</p><p dir="rtl">Yes</p></td></tr><tr><td><p>25.0)</p></td><td><p>24.6</p></td><td><p>±52.41</p></td><td><p>±52.71</p></td><td><p>69.</p></td><td><p>± 82.49</p></td><td><p>804.7</p></td><td><p dir="rtl">Neighborhood</p></td></tr><tr><td><p>0</p></td><td><p dir="rtl">٥/٥</p></td><td><p>0.50)</p></td><td><p>0).380</p></td><td><p>857</p></td><td><p>0.380</p></td><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">The closest</p><p>kl</p></td></tr><tr><td><p>23.5</p></td><td><p>27.1</p></td><td><p>±54.81</p></td><td><p>±52.17</p></td><td><p>71.</p></td><td><p>± 82.59</p></td><td><p>86.7</p></td><td><p dir="rtl">Museum</p></td></tr><tr><td><p>0</p></td><td><p dir="rtl">٥/٥</p></td><td><p>1.87</p></td><td><p>1.67</p></td><td><p>61</p></td><td><p>0.93</p></td><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Support</p></td></tr><tr><td><p>33.</p></td><td><p>12.9</p></td><td><p>±59.28</p></td><td><p>±77.89</p></td><td><p>89.</p></td><td><p>± 103.40</p></td><td><p>806.5</p></td><td><p dir="rtl">Packing</p></td></tr><tr><td><p dir="rtl">'h</p></td><td><p dir="rtl">٥/٥</p></td><td><p>0.54</p></td><td><p>0.71</p></td><td><p>44</p></td><td><p>0.53</p></td><td><p dir="rtl">٥/٥</p></td><td></td></tr><tr><td><p dir="rtl">Percentage of envy</p><p dir="rtl">N</p></td><td><p dir="rtl">Percentage of envy</p><p dir="rtl">N</p><p dir="rtl">{and</p></td><td><p>^-1-.^.-.1-111</p></td><td><p dir="rtl">11 for the lam {0e</p></td><td><p dir="rtl">The mistake was made</p><p dir="rtl">powerless</p></td><td></td><td><p dir="rtl">acceptance</p><p dir="rtl">Error</p><p dir="rtl">Actual (%)</p></td><td><p dir="rtl">the forest</p><p dir="rtl">Random</p><p dir="rtl">Yes</p></td></tr><tr><td><p>30.9</p></td><td><p>26.6</p></td><td><p>±56.87</p></td><td><p>±60.42</p></td><td><p>82.</p></td><td><p>±93.98</p></td><td><p>87.6</p></td><td><p dir="rtl">The decline</p></td></tr><tr><td><p>%</p></td><td><p dir="rtl">٥/٥</p></td><td><p>0.52</p></td><td><p>0.40</p></td><td><p>33</p></td><td><p>0.39</p></td><td><p>%</p></td><td><p dir="rtl">linear</p></td></tr><tr><td><p>30.8</p></td><td><p>28.4</p></td><td><p>±57.46</p></td><td><p>±59.48</p></td><td><p>83.</p></td><td><p>±95.08</p></td><td><p>87.3</p></td><td><p>LAS</p></td></tr><tr><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">٥/٥</p></td><td><p>0.52</p></td><td><p>0).380</p></td><td><p>0)0)</p></td><td><p>0).47</p></td><td><p dir="rtl">٥/٥</p></td><td><p>SO</p></td></tr><tr><td><p>22.9</p></td><td><p>13.2</p></td><td><p>±65.38</p></td><td><p>±73.63</p></td><td><p>804.</p></td><td><p>This is Q7.fi«</p></td><td><p>92.1</p></td><td><p dir="rtl">tree</p></td></tr><tr><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">٥/٥</p></td><td><p>0.47</p></td><td><p>0).54</p></td><td><p>81</p></td><td><p>0.53</p></td><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">The decision</p><p dir="rtl">to</p></td></tr><tr><td><p>13.3</p></td><td><p>14.3</p></td><td><p>±62.07</p></td><td><p>±61.41</p></td><td><p>71.</p></td><td><p dir="rtl">^-٦٦^</p></td><td><p>992.3</p></td><td><p dir="rtl">the forest</p></td></tr>
180476
-89-
<tr><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">٥/٥</p></td><td><p>0.41</p></td><td><p>0).31</p></td><td><p>62</p></td><td><p>0.53</p></td><td><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Random</p><p dir="rtl">Yes</p></td><td rowspan="4"></td></tr><tr><td><p>20.99</p><p>0</p></td><td><p dir="rtl">26.6</p><p dir="rtl">٥/٥</p></td><td><p>±49.83</p><p>0.52</p></td><td><p>±51.45</p><p>0).39</p></td><td><p>70.</p><p>0)4</p></td><td><p>± 82.05</p><p>0).34</p></td><td><p dir="rtl">804.9</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Nearest Neighbor</p><p>kl</p></td></tr><tr><td><p>23.5</p><p>0</p></td><td><p dir="rtl">1.5</p><p dir="rtl">٥/٥</p></td><td><p>±55.31</p><p>1.60</p></td><td><p>±51.41</p><p>1.33</p></td><td><p>1.</p><p>29</p></td><td><p>±83.38</p><p>1.32</p></td><td><p dir="rtl">86.7</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Museum</p><p dir="rtl">Support</p></td></tr><tr><td><p dir="rtl">37.8</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">25.3</p><p dir="rtl">٥/٥</p></td><td><p>±54.50</p><p>0.52</p></td><td><p>±65.51</p><p>1.06</p></td><td><p>87.</p><p>64</p></td><td><p>± 103.96</p><p>0.59</p></td><td><p dir="rtl">84.3</p><p dir="rtl">٥/٥</p></td><td><p dir="rtl">Packing</p></td></tr><tr><td colspan="9"><p dir="rtl">Also as expected, based on Figure 13, Table 5 indicates a significant improvement in the predictions.</p></td></tr>
DOBR selected from the reference model values for both baggage and random forest classifiers, see Figures 4a and 14b, respectively, below. Model (««34) (02) shows most of the improvements that suggest removing outliers before learning the model, along with DOBR classifier providing better results.
5 From simply using the DOBR classification of the full model (other than 9088). The difference in optimization results between the models shows that the choice of model is important. Although this decision is made by the analyst, it is interesting to compare the prediction accuracy of the model. The running time of the model and many other factors are also important and this research is not designed or intended to suggest the feasibility of one model over another.
The conclusion from Table 5 is clear and statistically significant. Given the potential for outlier bias, 10 as shown in the plot similar to Figure 13, a machine learning model using
DOBR methodology has better predictive accuracy for non-anomalous records than a machine learning model without DOBR.
Thus, an innovative illustrative computer system that includes a machine learning model with DOBR improves accuracy and reduces error in making predictions, thereby increasing the performance and efficiency of model execution. But the improvement can come at a price: there may be no predictive value or consideration for specific outliers. In models, 15 how outliers are modeled can vary based on the application.
180476
-90-
Table 6 shows the results of the predictive accuracy of the training/test sampling of the concrete compressive strength dataset using the packing classifier. The random forest classifier was not applied to this dataset. The table shows the root mean square error (see Eq. 15) at 95% confidence level between the test data and each model for 100 random selections of the training and test datasets.
5 Table 6
<tr><td><p>191, [3</p></td><td><p dir="rtl">Sorry, nooooo</p></td><td></td><td><p>!,fv, ,ΐ.fv٠„٠iii</p></td><td></td></tr><tr><td><p>0.08 ±9.58</p></td><td><p>0).085+10.13</p></td><td><p>0.1 ±9.2</p></td><td><p>0.07 ±10.49</p></td><td><p dir="rtl">For linear regression</p></td></tr><tr><td><p>0.09 ±9.53</p></td><td><p>0.08 ±10.25</p></td><td><p>0.1 ±9.2</p></td><td><p>0.07 ±10.51</p></td><td><p>LASSO</p></td></tr><tr><td><p>0.09 ±7.98</p></td><td><p>0.11 ±7.84</p></td><td><p>0.1 ±7.9</p></td><td><p>0.10 ±7.89</p></td><td><p dir="rtl">tree</p><p dir="rtl">For decisions</p></td></tr><tr><td><p>0.09 ±9.40</p></td><td><p>0.12 ±9.26</p></td><td><p>0.10 ±9.04</p></td><td><p>0.10 ±9.04</p></td><td><p dir="rtl">For the forest</p><p dir="rtl">Randomness</p></td></tr><tr><td><p>0.11 ±9.83</p></td><td><p>0.15 ±9.06</p></td><td><p>0.1 ±9.6</p></td><td><p>0.10 ±9.26</p></td><td><p dir="rtl">For the neighborhood</p><p dir="rtl">The closest to</p></td></tr><tr><td><p>0.11 ±10.32</p></td><td><p>0.15+10.09</p></td><td><p>0.2 ±10.6</p></td><td><p>0.11 ±9.84</p></td><td><p dir="rtl">support vector</p></td></tr><tr><td><p>0.11+9.44</p></td><td><p>0.12 ±8.82</p></td><td><p>0.1+9.3</p></td><td><p>0.10 ±9.02</p></td><td><p dir="rtl">To fill</p></td></tr>
The linear regression LASSOj produces the largest errors of the baseline or reference model. However, it produces
The models (1) have statistically similar prediction accuracy for all other models except the decision tree. In this case, the decision tree model produces the best prediction accuracy and all models except the linear regression LASSOj do not appear to improve with the addition of DOBR.
180476
-91-
Table 4 shows that there is little, if any, predictive improvement using the DOBR records chosen. This result is not surprising and is in fact expected based on the shape of the error tolerance versus model error curve shown in Figure 12.
Table 7 shows the increase (+) or decrease (-) in the prediction accuracy of the DOBR models relative to the model.
5 The benchmark in each case is, for example, the performance of the concrete compressive strength prediction accuracy of DOBR models: the packing classifier.
Table 7
<tr><td><p dir="rtl">percentage</p><p dir="rtl">To improve {y3}</p></td><td><p dir="rtl">percentage improvement</p><p>{y2}</p></td><td><p dir="rtl">percentage improvement</p><p>{y1}</p></td><td></td></tr><tr><td><p>± %8.63</p><p>%0.47</p></td><td><p>%0.51 ± %3.44</p></td><td><p>± %12.39</p><p>%0.55</p></td><td><p dir="rtl">Linear regression</p></td></tr><tr><td><p>± %9.29</p><p>%0.42</p></td><td><p>%0.53 ± %2.44</p></td><td><p>± %11.98</p><p>%0.60</p></td><td><p>LASSO</p></td></tr><tr><td><p>± %1.28-</p><p>%0.77</p></td><td><p>± %2.54-</p><p>%1.29</p></td><td><p>± %0.44-</p><p>%1.07</p></td><td><p dir="rtl">tree of decision</p></td></tr><tr><td><p>± %6.17-</p><p>%0.67</p></td><td><p>± %3.68-</p><p>%0.41</p></td><td><p>± %6.73-</p><p>%1.46</p></td><td><p dir="rtl">Random Forest</p></td></tr><tr><td><p>± %0.23</p><p>%0.19</p></td><td><p>%0.66 ± %2.11</p></td><td><p>± %4.17-</p><p>%0.99</p></td><td><p dir="rtl">The neighborhood is closer to k</p></td></tr><tr><td><p>± %4.88-</p><p>%0.29</p></td><td><p>± %2.61-</p><p>%1.38</p></td><td><p>± %7.38-</p><p>%1.37</p></td><td><p dir="rtl">support vector</p></td></tr>
18476
-92-
<tr><td><p>± %4.77</p><p>%0.98</p></td><td><p>%1.20 ± %2.11</p></td><td><p>± 02.71-</p><p>%1.17</p></td><td><p dir="rtl">Packing</p></td></tr>
These results are not surprising because the model error curves versus the error acceptance curves for the LASSOj linear regression were the ones with the largest nonlinearities and the others were almost straight lines indicating that models that adequately predict the target variable and anomalous analysis are not required. This is the message in Table 7. The model outputs with respect to the predicted concrete compressive strength 5 are represented in Appendix A attached.
Now looking at the power consumption prediction error results in Table 8, there is a different situation involving, for example, the device power consumption prediction errors of the bagging and random forest classifiers. The bagging and 5_l_ASSO linear regression models have the largest reference prediction errors and the random forest model has the smallest. The DOBR model errors in the right three columns show that in many cases, the DOBR models produce higher prediction accuracy than the reference models.
Table 8
<tr><td></td><td><p dir="rtl">1.31 a</p></td><td></td><td><p>!,fv, ,ΐ.fv٠„٠iii</p></td><td><p dir="rtl">Packing</p></td></tr><tr><td><p>0.47 ±92.36</p></td><td><p>0.32 ±86.47</p></td><td><p>0.39 ±84.70</p></td><td><p>0.41 ±93.99</p></td><td><p dir="rtl">For linear regression</p></td></tr><tr><td><p>0.44 + 94.06</p></td><td><p>0.32 ±85.76</p></td><td><p>0.39 ±84.87</p></td><td><p>0.37 ±94.87</p></td><td><p>LASSO</p></td></tr><tr><td><p>0.52 ±86.37</p></td><td><p>0.49+93.34</p></td><td><p>0.54 +87.804</p></td><td><p>0.51 ±92.06</p></td><td><p dir="rtl">Decision tree</p></td></tr><tr><td><p>0.41 ±79.08</p></td><td><p>0.35+80.57</p></td><td><p>0.39 ±81.82</p></td><td><p>0.36+11.13</p></td><td><p dir="rtl">For the forest</p><p dir="rtl">Randomness</p></td></tr><tr><td><p>0.45 ±82.31</p></td><td><p>0.33 ±84.92</p></td><td><p>0.38 + 84.75</p></td><td><p>0.38 ±82.49</p></td><td><p dir="rtl">The closest neighbor</p><p dir="rtl">A</p></td></tr>
180476
-93-
<tr><td><p>1.08 ±84.29</p></td><td><p>1.10+77.46</p></td><td><p>1.20+79.27</p></td><td><p>0.93 ±82.59</p></td><td><p dir="rtl">support vector</p></td></tr><tr><td><p>0.58 ±97.40</p></td><td><p>0.71+92.52</p></td><td><p>0.46+85.55</p></td><td><p>0.53+103.40</p></td><td><p dir="rtl">To fill</p></td></tr>
<tr><td><p>191, [3</p></td><td></td><td><p dir="rtl">Spelling</p></td><td><p>!,fv, ,ΐ.fv٠„٠iii</p></td><td><p dir="rtl">Random Forest</p></td></tr><tr><td><p>0.45 ±91.75</p></td><td><p>0.33+84.45</p></td><td><p>0.40 ±81.95</p></td><td><p>0.39+93.98</p></td><td><p dir="rtl">Linear regression</p></td></tr><tr><td><p>0.54+93.84</p></td><td><p>0.38+83.93</p></td><td><p>0.46 ±82.53</p></td><td><p>0.47 ±95.08</p></td><td><p dir="rtl">Yes</p></td></tr><tr><td><p>0.45 ±85.71</p></td><td><p>0.49+93.34</p></td><td><p>0.46 ±87.11</p></td><td><p>0.53 ±92.08</p></td><td><p dir="rtl">Decision tree</p></td></tr><tr><td><p>0.37 ±78.14</p></td><td><p>0.35+ 78.92</p></td><td><p>0.37+97.34</p></td><td><p>0.35+77.59</p></td><td><p dir="rtl">For the random forest</p></td></tr><tr><td><p>0.39+81.51</p></td><td><p>0.27+83.59</p></td><td><p>0.31 ±82.62</p></td><td><p>0.34 ±82.50</p></td><td><p dir="rtl">The closest neighbor to</p><p>k</p></td></tr><tr><td><p>1.42 ±85.24</p></td><td><p dir="rtl">66.66 B D4M1</p></td><td><p>1.55 + 79.76</p></td><td><p>1.32 ±83.38</p></td><td><p dir="rtl">support vector</p></td></tr><tr><td><p>0.59+97.76</p></td><td><p>0.79+93.55</p></td><td><p>0.51 ±85.94</p></td><td><p>0.59+103.96</p></td><td><p dir="rtl">To fill</p></td></tr>
It is interesting to note that the reference packing model has the largest reference error values, but the results
The augmented DOBR model is generally in the same statistical ranges as the other models.
Additionally, for practical reasons, the support vector model was run for only 10 iterations. This explains why
5 Increased uncertainty across the results of its model.
Detailed optimization results are shown in Table 9 relating, for example, to the performance of the device power consumption prediction accuracy of the DOBR models. Note that at least one of the DOBR models produces
Increase in prediction accuracy for most machine learning models. However, there are also relatively large differences so there are no conclusive results regarding the improvement in predictive ability resulting from DOBR. From the model error curves 10 vs. error acceptance for the power data, all plots show nonlinear behavior with forest models
180476
-94-
Randomness and decision trees with the smallest curvature. Models, particularly random forest, appear to be able to model this variation adequately based on the results shown here. The model outputs in terms of predicted energy use are represented in Appendix B attached.
Table 9
<tr><td><p dir="rtl">The percentage of improvement</p><p dir="rtl">[3</p></td><td><p dir="rtl">Percentage improvement 0</p><p dir="rtl">[2</p></td><td><p dir="rtl">percentage of improvement</p></td><td><p dir="rtl">Packing</p></td></tr><tr><td><p dir="rtl">%1٠74 ± %0٠11</p></td><td><p dir="rtl">٥/٥7.98 ± %0.23</p></td><td><p>%0.27+%9.87</p></td><td><p dir="rtl">linear regression</p></td></tr><tr><td><p dir="rtl">%0.85 ± %0٠10</p></td><td><p>%0021 ± 5/ο9.59</p></td><td><p>%0.23 ± %10.53</p></td><td><p>LASSO</p></td></tr><tr><td><p dir="rtl">%6.16 ± %0٠45</p></td><td><p>00.36+%1.41-</p></td><td><p>%0.56 ± %4.55</p></td><td><p dir="rtl">Decision tree</p></td></tr><tr><td><p>5/ο0033 ± 5/ο1.74</p></td><td><p>00.41+%3.68-</p></td><td><p>00.41* %5.28-</p></td><td><p dir="rtl">For the random forest</p></td></tr><tr><td><p>%0.0019 ± 5/ο0.23</p></td><td><p>5/ο0029 ± 5/ο2.96-</p></td><td><p>00.30)+%2.74-</p></td><td><p dir="rtl">Nearest neighbor kJ</p></td></tr><tr><td><p>5/ο0024 ± 5/ο2005-</p></td><td><p dir="rtl">%0٠95±%6٠21</p></td><td><p dir="rtl">%4٠02 ± %0٠93</p></td><td><p dir="rtl">support vector</p></td></tr><tr><td><p>00.48+%5.77</p></td><td><p>5/ο0075 ± 10048%</p></td><td><p dir="rtl">%0٠42±%17.23</p></td><td><p dir="rtl">To prepare</p></td></tr>
<tr><td><p dir="rtl">The percentage of improvement</p><p>y3}</p></td><td><p dir="rtl">percentage of improvement</p><p dir="rtl">[2</p></td><td><p dir="rtl">Percentage of envy</p><p>{yi}</p></td><td><p dir="rtl">Random Forest</p></td></tr><tr><td><p>%0.12 ± %2.38</p></td><td><p>%0.23 ± %10.14</p></td><td><p>%0.27 ± %12.80</p></td><td><p dir="rtl">linear regression</p></td></tr><tr><td><p>%0.12±%1.31</p></td><td><p>%0.21± %11.71</p></td><td><p>%0.24± %13.20</p></td><td><p dir="rtl">a)15</p></td></tr><tr><td><p>%0.41±%6.89</p></td><td><p>%0.44± %1.40-</p></td><td><p dir="rtl">%0٠58±%5٠35</p></td><td><p dir="rtl">Decision tree</p></td></tr>
180476
-95-
<tr><td><p>00.36+%0.73-</p></td><td><p>%0.39± %1.74-</p></td><td><p>%0.44 ± %2.28-</p></td><td><p dir="rtl">For the random forest</p></td></tr><tr><td><p>%0.16 ± %1.20</p></td><td><p>%0.30± %1.34-</p></td><td><p>%0.32± %0.16-</p></td><td><p dir="rtl">Nearest neighbor d k</p></td></tr><tr><td><p>%0.27 ± 5/ο2.23-</p></td><td><p dir="rtl">%0٠90±%6٠73</p></td><td><p dir="rtl">%0.98±%4٠35</p></td><td><p dir="rtl">support vector</p></td></tr><tr><td><p>%0.48 ± %5.94</p></td><td><p>%0.77+%9.98</p></td><td><p>%0.47 ± %17.31</p></td><td><p dir="rtl">To fill</p></td></tr>
Figures 14a and 14b show plots of non-abnormal and abnormal distributions in the classifier models according to
An illustrative embodiment of an inventive computer-based illustrative system having a DOBR classifier according to one or more embodiments of the present disclosure.
The concrete dataset is relatively small, so data charts can provide visual insights,
<p dir="rtl">5 But since DOBR has little value in this case, plotting this dataset does not improve our understanding of how DOBR works. However, for the energy dataset predictions, DOBR does produce some significant predictive improvements. But its relatively large size (13,814 training records, 5,921 test records) makes it difficult to interpret straightforward scatter plot visualizations. Scatter plots, such as Figure 9 and Figure 10, with a large number of points can obscure any detail. The 10-error optimization results shown in Table 3 are summaries of non-outlier datasets, but the question remains as to how the DOBR method and classification model produce these results.</p>
In models, to address this question, the error distributions of two representations of the model can be analyzed: the random forest classifier (Figure 14a) and the bagging classifier (Figure 14b) for both anomalous and non-anomalous data sets. In one model, the non-anomalous errors should be smaller than the
15 Anomalous by design but DOBR's innovative explanatory model and classification process are built exclusively from training data so that the test dataset can contain information that has never been seen before. Therefore, the model and classification calculations may not be accurate and can be visualized.
180476
-96-
Classification errors in these schemes. This work is performed for linear regression models and bagging regression models as these approaches have the largest and smallest improvement benefits, respectively, as shown in Tables 5.
For discussion, the reference error value is marked in both the plots in Figure 14a and Figure 14b. The top set of arrows shows that 80% of the non-abnormal error values are less than 1000 which assumes
<p dir="rtl">5 That 20% of the error values are < 1000. This lower set of arrows also shows that for skewed distributions, about 20% of the outliers have an error > 1000 or 80% have errors < 1000 - which should represent outlier errors. Without knowing the percentage of error tolerances, we cannot calculate the accuracy of the classification process precisely but the above plots indicate that although misclassification occurs, most values are correctly classified.</p>
<p dir="rtl">10 Figure 14C shows model error plots as a function of error tolerance values for an illustrative use case of an illustrative model of an innovative computer-based illustrative system with a DOBR-trained machine learning model for predicting the off-time of a well drilling in accordance with one or more models of the present disclosure.</p>
Offshore drilling operations present unique challenges for the oil and gas industries. In addition to
<p dir="rtl">15 Logistical and environmental risks that can be observed from weather and ocean depths, there are hidden downhole risks operating in high temperature, pressure and vibration environments. Drilling times are set to tight schedules and delays due to downhole equipment failure (non-productive time or NPT) can represent significant revenue losses.</p>
To assist in NPT management, a machine learning model is being built to help predict future downtime events20 for the purpose of incorporating these estimated delays into the contract terms that specify drilling targets. Historical event monitoring includes: distance drilled [feet], hole size [inches], tool size [inches], site pressure intensity, maximum slope [degrees/100 feet], vibration intensity class, curvature class, and NPT (hours).
18476
-97-
Linear, xgboost, gradient boosting, and random forest regression models were applied to downhole equipment failure data with a 20/80 training/test split to measure the predictive accuracy of the model. Hyperband was used to tune the models and the relevant parameter values are shown in Table 10 below:
Table 10
<tr><td><p>eta = 0.76, max_depth = 4, min_child_weight = 0.43</p></td><td><p>xgboost</p></td></tr><tr><td><p>learning_rate = 0.34, min_samples_split = 0.58, n_estimators = 13</p></td><td><p dir="rtl">progressive reinforcement</p></td></tr><tr><td><p>max_depth = 4, min_samples_leaf = 2, min_samples_split = 9, n_estimators = 6</p></td><td><p dir="rtl">Random Forest</p></td></tr>
<p dir="rtl">5 The classification function that transfers DOBR-computed outlier information to the test dataset can be chosen as a random forest model with a number of estimation units equal to, say, 5. This tuning activity is also performed in the training part of the analysis. The parameter selection metric is to calculate the percentage of correctly classified items for the training set and compare it to the model's error tolerance value.</p>
10 Linear regression is included in this analysis because it is the only model where the coefficients can provide engineering insights to help identify additional best practice improvements. Other models are more powerful from a predictive perspective but provide little insight.
As discussed in this specification, there are several DOBR-related models that can be built upon a basic DOBR process. In this example, three models are presented: M
15 It represents a specific model that has been highly tuned.
Using typical values and outliers defined by DOBR for training and test datasets:
fake code 8
18476
-98-
[038-4001= = 01 ([0 -51, [0-5031)
0030+500+95+1 50
Dad6•5G0T4•4••,•361AD•-405614-•A••01 = 50 10101 [08•
0030191 03 90+1
Ba.seM٠d.el_^ta١, S 5βΜ٠ά61_^٠٩.Λ = m540 what
where 5271-+1-7 and 5271-+1-7 are typical values computed by DOBR from the LETRIP set, 53-**-a and 53-*t-a are outliers computed by DOBR from the training set, 250-/1-0052 are typical values and outliers for the test dataset, respectively, computed from the 5-DOBR classification model, and 7252140-12) BaseModeLyout are model results computed by non-
DOBR is classified into typical values and outliers using the DOBR classification model, and I sets the BaseModel values to 1-3852402 for typical values defined by DOBR and to 8856060 for outliers defined by DOBR.
Of these subgroups, the three DOBR models are:
A 038, 444-45 1031-9088 = 1001-708
B 10821-304540081,5038 - 100212-9988
J 9832981 149813-0988 = 4983-903
Running percentage error tolerance curves versus model error curves for models
The above-mentioned fine-tuning produces the curves shown in Figure 14c. The important property of these curves is their curvature—not the error values themselves. In general, the more
The steeper the slope of the linear curve over the range (0, 100%), the less influence the outliers have. For offshore downhole equipment failure data, the curves show linearity up to about 80% error tolerance.
180476
-99-
Then various nonlinear tendencies appear. When analyzing the curve as a function of error tolerance values, the following table (Table 11) shows the error tolerance threshold values specified for DOBR analysis.
Table 11
<tr><td><p dir="rtl">Percentage of error acceptance applied</p></td><td><p dir="rtl">Regression model</p></td></tr><tr><td><p>85.0</p></td><td><p dir="rtl">linear</p></td></tr><tr><td><p>85.0</p></td><td><p>xgboost</p></td></tr><tr><td><p>85.0</p></td><td><p dir="rtl">progressive reinforcement</p></td></tr><tr><td><p>85.0</p></td><td><p dir="rtl">Common Forest</p></td></tr>
All models were run using calculated hyperparameters and assigned error tolerance values. The following are represented:
5 The model outputs for the predicted NPT are shown in Appendix C, attached herewith, and the error results are shown in Table 12 below:
Table 12
<tr><td><p dir="rtl">model</p><p>DOBR number</p><p>3</p></td><td><p dir="rtl">model</p><p>DOBR number</p><p>2</p></td><td><p dir="rtl">model</p><p>DOBR number</p><p>1</p></td><td><p dir="rtl">The error is fundamental.</p><p>(DOBR without)</p></td><td><p dir="rtl">Regression model</p></td></tr><tr><td><p>15.6</p></td><td><p>14.9</p></td><td><p>14.0</p></td><td><p>16.4</p></td><td><p dir="rtl">linear</p></td></tr><tr><td><p>11.6</p></td><td><p>10.0</p></td><td><p>10.6</p></td><td><p>11.1</p></td><td><p>xgboost</p></td></tr><tr><td><p>9.6</p></td><td><p>17.8</p></td><td><p>10.5</p></td><td><p>16.9</p></td><td><p dir="rtl">progressive reinforcement</p></td></tr><tr><td><p>13.4</p></td><td><p>9.0</p></td><td><p>9.0</p></td><td><p>13.9</p></td><td><p dir="rtl">Common Forest</p></td></tr>
18476
-100-
Now that we have a non-DOBR model along with the three DOBR models, we are in a position to choose which model to use in production for future forecasts. Generally, the linear model provides the least predictive accuracy and DOBR models #1 or #2 provide the best predictive accuracy. At this point, the analyst can balance these accuracy figures with other practical considerations, such as
<p dir="rtl">5 Computing time to define a model to apply to future predictions.</p>
While the results are specific to the use of DOBR for training and implementing machine learning models for application in concrete compressive stress prediction and energy prediction, other applications are also considered.
For example, image display and visualization can benefit from machine learning models to predict and automatically implement display parameters based on, for example, medical data, as shown in
<p dir="rtl">10 US Patent No. 10339695, incorporated herein in its entirety for reference to all</p>
Objectives. DOBR can be used to train and execute machine learning models for content-based rendering. A medical dataset representing a 3D region of a patient can be used as input data. Using DOBR, outliers can be removed from a medical training dataset so that a machine learning model can be trained on non-outlier data according to the DOBR techniques described above. The model is trained
<p dir="rtl">15 Machine learning on deep learning of non-anomalous data from a training medical dataset to extract features from the medical dataset and to output values for two or more physical display parameters based on the input of the medical dataset. In some embodiments, the physical display parameters are controllers for appropriate data processing, lighting design, display design, material suitability, or an internal display property. The display module provides</p>
<p dir="rtl">20 On a physical basis, a realistic 3D image of the patient's area is generated using the output values generated by the application.</p>
In another illustrative application of DOBR for training and executing machine learning models, a machine learning model may be trained using the DOBR techniques described above to generate a control command for a machine to output the control command, as described in U.S. Pat. 10,317,854, incorporated herein in its entirety by reference.
<p dir="rtl">25 For all purposes. In such an example, the simulator might simulate the operation of the machine based on</p>
18476
-101-
Control command. The simulator may generate a complete dataset to train a machine learning model by simulating the physical actions of the machine based on the control command. This dataset may be processed using DOBR algorithms to ensure that any simulation anomalies are removed when training the model parameters including the work process data, control command data, and machine data used as input to each simulation.
<p dir="rtl">5 Other examples of DOBR implementations for training and executing machine learning models could include, for example, Software-as-a-Service implementations for on-demand training and deployment, outlier dataset analytics with outlier trained models, network energy optimization modeling, and user content recommendation modeling to improve user engagement, among other implementations. Some examples are described in more detail below:</p>
<p dir="rtl">10 SaaS implementation for custom ML training</p>
Figure 15 illustrates a framework diagram of an explanatory bias-reduced model generation service for training and deploying a machine learning model according to one or more of the embodiments of the present disclosure.
In some embodiments, the bias-reduced model generation service 1500 may be implemented as
SaaS (Software-as-a-Service) by including bias reduction components
<p dir="rtl">15 Dynamic Outlier Reduction (DOBR) is used in training datasets to train and deploy one or more machine learning models. In some embodiments, DOBR provides an iterative process to remove outliers subject to a predefined criterion. This criterion is a user-defined error tolerance value expressed as a percentage. This criterion indicates the amount of error that the user is willing to accept in the model based on their insights and other analysis results described later in this discussion. The value indicates</p>
<p dir="rtl">20 100% to accept all errors and no records will be removed in the DOBR process. If selected,</p>
0%, all records will be deleted. In general, error tolerance values in the range of 80 to 95% have been observed for industrial applications.
In some embodiments, the user may interact with the bias-reduced model generation service 1500 to initiate a machine learning model request. In some embodiments, the bias-reduced model generation service 1500 may receive
18476
-102-
Request, train machine learning models on request and return a machine learning model to the user for use in the user's purpose.
In some embodiments, the user may use a computing device 1511 to communicate with the bias-reduced model generation service 1500, for example, over a network 1520. In some embodiments, the user may
<p dir="rtl">5 The computing device 1511 sends a model request 1512 to the bias-reduced model generation service 1500 to request a custom trained model. Accordingly, the model request 1512 may include the required model attributes such as the error tolerance value for filtering the DOBR dataset, the model type (e.g., classification, object detection, natural language processing, data prediction, time series prediction, computer vision, etc.), the model memory limits, or any other illustrative model attributes.</p>
<p dir="rtl">10 Other required or any combination thereof. In some embodiments, the request for Form 1512 may also include training data for the modeling task for which the trained custom model will be used. For example, for a grid energy optimization model, the request may contain a package of electric power demand data according to, for example, time of day, day of week, day of month, month of year, season, weather, location, population density, among other energy demand data</p>
<p dir="rtl">15 Other electrical. For example, for content recommendation and advertising models for browsing online content for one or more users, the request may contain a package of user interaction data including, for example, click-through rates, click-through frequency, times spent on the content, location of the content on a page, content screen area, content type or category, among other user interaction data, in combination with user data such as user characteristics, including</p>
<p dir="rtl">20 This may include, for example, browser, location, age, or other user characteristics or any combination thereof.</p>
In some embodiments, the computing device 1511 may send the model request 1512 containing the training data to the low-bias model generation service 1500 over the network 1520 using any appropriate electronic request. In some embodiments, the model request 1512 may be connected to the model generation service
25 Low bias 1500 via, for example, a suitable application programming interface (API), protocol
18476
-103-
m transmitter, or other communication technology. In some embodiments, the form request 1512 may be communicated via, for example, a direct interface between the computing device 1511 and the low-bias form generation service 1500 or via the network 1520 (such as a local area network (LAN), wide area network (WAN), Internet, intranet, or other network and combinations thereof), or combination thereof. In
<p dir="rtl">5 In some embodiments, the connectivity may include, for example, hard wired connections (e.g., fiber optic cables, coaxial cables, copper wire cables, Ethernet, etc.), wireless connections (e.g., Z-Wave, Zigbee, Bluetooth, WiFi, cellular networks such as 4G, 5G, Long Term Evolution (LTE), High Speed Downlink Packet Access (HSPA), Global System for Mobile Communications (GSM),</p>
<p dir="rtl">10 Code Division Multiplexing (CDMA) or other technologies, and combinations thereof, or a combination thereof.</p>
Manage the error tolerance value via a user input device 1508 and view the results via a display 1512, among other user interaction behaviors using the display 1512 and the user input device 1508. Based on the error tolerance value, the low-bias model generation service 1500 may analyze a dataset 1511 received into a database 1510 or other
<p dir="rtl">15 of storage units that are in contact with the low-bias model generation service 1500. The low-bias model generation service 1500 may receive the dataset 1511 via the database 1510 or other storage device and make predictions using one or more machine learning models that have low dynamic anomaly bias to improve accuracy and efficiency.</p>
In some embodiments, the low-bias model generation service 1500 may include a combination of 20 hardware and software components, including, for example, storage devices and memory, storage memory
Temporary, buffers, bus, input/output (I/O) interfaces, processors, controllers, networking and communications devices, operating system, kernel, and device drivers, among other components. In some embodiments, the processor 1507 is in communication with multiple other components to perform the functions of the other components. In some embodiments, each component has a scheduled time on the processor 1507 to perform the functions of the component, but in some embodiments, each component is scheduled to a processor
18476
-104-
One or more processors in a processing system 1507. In other embodiments, each component has its own processor embedded therein.
In some embodiments, the low-bias model generation service components 1500 may include, for example, a DOBR training engine 1501 in communication with a model index 1502, a library
<p dir="rtl">5 Illustrative 1503, regression variable parameter library 1505, classifier parameter library 1504, and DOBR filter 1506, among other possible components. Each component may contain a combination of hardware and software to implement the functions of the components, such as memory and storage devices, processing devices, communications devices, input/output (I/O) interfaces, control units, networking and communications devices, an operating system, a kernel, device drivers, and a set of instructions, among other components.</p>
<p dir="rtl">10 In some embodiments, the DOBR training engine 1501 includes a model engine for instantiating and executing machine learning models. The DOBR training engine 1501 may access models for instantiation in the model library 1503 by using the model index 1502. For example, the model library 1503 may contain a library of machine learning models that may be selectively accessed and instantiated for use by an engine such as the DOBR training engine 1501. In some</p>
<p dir="rtl">15 Models The model library 1503 may contain machine learning models such as a support vector machine (SVM), a linear regression variable, a Lasso model, a decision tree regression variable, decision tree classifiers, a random forest regression variable, random forest classifiers, K-neighbor regression variables, K-neighbor classifiers, gradient boosting regression variables, gradient boosting classifiers, among other possible classifiers and regression variables.</p>
<p dir="rtl">20 In some embodiments, based on the characteristics of the model request form 1512, the DOBR training engine 1501 may select a set of model structures. For example, some models may be smaller than others, and then based on the size requirements in the model request form 1512, the DOBR training engine 1501 may use the model index 1502 to identify model structures in the model library 1503 that meet the maximum size requirements. Similarly, a model type or</p>
<p dir="rtl">25 A task type to specify model structures. For example, a DOBR training engine might choose</p>
18476
-105-
1501 A collection of model architectures included for use with classification tasks, regression tasks, time series forecasting tasks, computer vision tasks, or any other task.
Accordingly, in some embodiments, to facilitate access to the machine learning model library in the model library 1503, the DOBR training engine 1501 may use the model index 1502. In
<p dir="rtl">5 Some models, the model index 1502 may index each model by reference to a model identifier, model type, set of task types, memory footprint, among other model architecture properties. For example, models that include, for example, linear regression, XGBoost regression, support vector regression, Lasso, K-neighbor regression, bagging regression, gradient boosting regression, random forest regression, decision tree regression, among other regression models and 10 classification models may be indexed by a numerical identifier and labeled with a specific name.</p>
In some embodiments, the program instructions are stored in the memory of the relevant model library 1503 or model index 1502 and are temporarily stored in a cache for availability to the processor 1507. In some embodiments, the DOBR training engine 1501 may use the model index 1502 by accessing or communicating with the index via communications and/or input/output devices, wherein the index is 15 used to invoke models as functions from the model library 1503 via communications and/or input/output devices.
In some embodiments, to facilitate optimization and customization of models called by the DOBR training engine 1501, the low-bias model generation service 1500 may record model parameters in, for example, memory or storage, for example, hard disk drives.
<p dir="rtl">20 Solid state disk drives, random access memory (RAM), flash storage, among other storage and memory devices. For example, regression variable parameters may be recorded and modified in the regression variable parameter library 1505. The regression variable parameter library 1505 may then contain storage and communications devices configured with sufficient memory and bandwidth to store, modify, and communicate a large number of parameters for many regression variables, e.g.</p>
<p dir="rtl">25 Example, in real time. For example, for each machine learning classification model an instance is created.</p>
18476
-106-
By means of the DOBR training engine 1501, the specified parameters may be initialized and updated in the regression variable parameter library 1505. In some embodiments, the user may, by requesting the model 1512 from the computing device 1511, generate an initial set of parameters in addition to the training data. However, in some embodiments, the initial set of parameters may be predetermined or random 5 (e.g., randomly initialized). When creating an instance of a classification machine learning model, the user may associate
The DOBR training engine 1501 selects a model from the model index 1502 with a set of parameters in the regression variable parameter library 1505. For example, the DOBR training engine 1501 may call a set of parameters according to, for example, an identification number (ID) associated with a particular regression model.
<p dir="rtl">10 Similarly, in some embodiments, classifier parameters may be recorded and modified in the classifier parameter library 1504. The classifier parameter library 1504 may therefore include storage and communications devices configured with sufficient memory and bandwidth to store, modify, and communicate a large number of parameters for multiple classifiers, e.g., in real time. For example, for each classification machine learning model instantiated by the DOBR training engine 1501, the parameters may be initialized</p>
<p dir="rtl">15 specified and updated in the classifier parameter library 1504. In some embodiments, the user, via the user input device 1508, may create an initial set of parameters. However, in some embodiments, the initial set of parameters may be pre-specified or random (e.g., randomly initialized). When creating an instance of a machine learning model for classification, the DOBR training engine 1501 may associate a model selected from the model index 1502 with a set of parameters in the DOBR library.</p>
<p dir="rtl">20 Classifier parameter 1504. For example, a DOBR training engine 1501 may call a set of parameters according to, for example, an identification number (ID) associated with a particular regression model.</p>
In some embodiments, by invoking and receiving a set of models from the model library 1503 via the model index 1502 and relevant parameters from the regression variable parameter library 1505 and/or the classifier parameter library 1504, the DOBR training engine 1501 may load one or more
<p dir="rtl">25 From models that have been instantiated and initialized, for example, in a cache or module.</p>
18476
-107-
DOBR Training Engine 1501 Caching. In some embodiments, the training dataset may be ingested from the model request 1512, and the DOBR Training Engine 1501 may train each model in the model set using an iterative DOBR training procedure.
In some embodiments, for example, the processor 1507 or a processor in the drive may be used.
<p dir="rtl">5 Training 1501 DOBR Each model transforms the training dataset into, for example, a specific forecast, for example, a predicted demand for electrical energy for the grid based on each input data point, for example, time of day, day of week, day of month, month of year, season, weather, location, population density, among other electrical energy demand data. The predicted output can be compared to the actual energy demand for the training dataset.</p>
<p dir="rtl">10 Similarly, for example, the DOBR training engine 1501 may train the model set to model user interaction based on content attributes based on the training dataset for the model request 1512. For example, the model set may be used to predict a predicted user interaction based on inputs from the training dataset including, for example, the location of the content on the page, the content screen area, the content type or classification, among other interaction data.</p>
<p dir="rtl">15 Other user data, in conjunction with user data such as user characteristics, including but not limited to browser, location, age, or other user characteristics, or any combination thereof. The predicted user engagement can then be compared to actual user engagement for each input according to the training dataset based on user engagement metrics such as click-through rates, click-through frequency, time spent on content, among other user engagement metrics or any combination thereof.</p>
<p dir="rtl">20 However, in some embodiments, outliers in the training dataset can reduce the accuracy of the model 1512 request, thereby increasing the number of training iterations to achieve an accurate set of parameters for a given model in a given application. To improve accuracy and efficiency, the DOBR training engine 1501 may include a DOBR filter 1501 to dynamically test the errors of data points in the training dataset to identify outliers. The outliers can then be removed to provide a set of</p>
<p dir="rtl">25 Training data that is more accurate or representative than a Form 1512 request. In some embodiments, it may provide:</p>
18476
-108-
DOBR filter 1501b is a recursive mechanism for removing anomalous data points subject to a predetermined criterion, e.g., the user-defined error acceptance value described above and provided, e.g., by a user via the user input device 1508. In some embodiments, the user-defined error acceptance value is expressed as a percentage where, for example, a value of 100%5 indicates that all errors will be accepted and no data points will be removed by filter 1501b, whereas
A value of, for example, 0% removes all data points. In some embodiments, the filter 1501b may be configured with an error acceptance value in the range between, for example, about 80% and about 95%.
In some embodiments, the DOBR filter 1501 operates in conjunction with an optimizer 1506, which is configured to identify the error and optimize the parameters for each model in the regression variable parameter library 1505.
and the classifier parameter library 1504. Thus, in some embodiments, the optimizer 1506 may identify the model and communicate the error to the filter 1501b of the DOBR training engine 1501. Thus, in some embodiments, the optimizer 1506 may, for example, have storage and/or memory devices and communication devices with sufficient memory capacity and bandwidth to receive the data set 1511 and the model predictions and determine, for example, outliers, convergence, error, absolute value error, among other error metrics.
In some embodiments, the DOBR training engine 1501 selects and trains multiple models using the DOBR filter 1501 and the training dataset from the model request 1512, the DOBR training engine 1501 may compare the error rates between each model in the last iteration of training. The DOBR training engine 1501 may then examine each model for the lowest error rate using the set of
Low outlier data for the training dataset. The model with the lowest error can be considered the best performing model and thus can be selected for deployment. In some embodiments, the 1501 DOBR training engine may select a set of models that includes only one model. In such a scenario, the 1501 DOBR training engine may skip the step of comparing error rates and use the 25 single models for deployment.
18476
-109-
In some embodiments, to facilitate deployment, the low-bias model generation service 1500 may return the selected model, trained using the dynamic low-bias anomaly training filter 1501b, to the computing device 1511 as a production-ready model 1513. In some embodiments, the production-ready model 1513 may include the selected model architecture as requested by model 1512.
<p dir="rtl">5 and trained parameters of the model. Thus, in some embodiments, the Low-Bias Model Generation Service 1500 can provide a SaaS solution for on-demand training and deployment of machine learning models, specifically selected and trained for a production environment and/or a particular user task. Thus, the user can simply develop AI software products without having to build a machine learning model from scratch. Furthermore, using DOBR improves the accuracy and efficiency of model training by removing</p>
<p dir="rtl">10 Dynamically remove outliers from the training dataset to reduce bias and error in the model.</p>
Outlier Dataset Analysis
Figures 16A and 16B illustrate the reduction of anomalous dynamic bias for modeling an anomalous data set according to an illustrative methodology in accordance with one or more of the embodiments of the present disclosure.
In some embodiments, one or more models can be trained to predict an output based on a given input.
<p dir="rtl">15 1606 x. In some embodiments, DOBR, such as the DOBR training engine 1501 and the filter, provides</p>
1501b described above, is an iterative process for removing anomalous records subject to a predefined criterion. This condition is a user-specified error acceptance value expressed as a percentage. This condition indicates the amount of error the user is willing to accept in the model based on their insights and other analysis results described later in this discussion. A value of 100% indicates that all
<p dir="rtl">20 Errors and no records will be removed in the DOBR process. If 0% is selected, all records will be removed. In general, error tolerance values in the range of 80 to 95% have been observed for industrial applications.</p>
In some embodiments, as described, bias reduction through iterative and dynamic anomaly reduction in training machine learning models can provide more efficient and robust training of machine learning models.
18476
-110-
Accuracy. In some models, in addition to modeling a low-level anomaly dataset, machine learning models can be applied in addition to other analytical models to an anomaly dataset. This modeling of an anomaly dataset can lead to accurate knowledge of anomalies such as extreme events, external factors, outliers, and the root causes of such anomalies.
<p dir="rtl">5 In some embodiments, the DOBR anomaly analysis may include a pre-analysis where an error acceptance criterion (m) is specified, such as mm = 80%. In some embodiments, the error acceptance criterion (C) may be defined in accordance with, for example, Equation 1 as described above. In some embodiments, while other functional relationships may be used to adjust c(a), the percentile function is an intuitive guide to understanding why the model includes or excludes certain data records, such as Equation 2 as</p>
<p dir="rtl">10 As shown above. Since the DOBR procedure is iterative, in one embodiment, a convergence criterion of, for example, 0.5% can be specified.</p>
In one embodiment, given a dataset (0.1 1604 , a solution model 1608 M , and an error tolerance criterion 1624 OC ), DOBR may be implemented to reduce bias in training the model 1608 M . In some embodiments, the solution model 1608 M is implemented by a model engine, including, for example, a processing device, memory, and/or storage device. According to one embodiment, the explanatory methodology computes
Model coefficients, ()/1602 and model estimates 13 1610 for all records applying the solution model, 1608 M, to the full input data set (-7( according to, for example, Equation 3 as shown above.
Then, according to an illustrative embodiment, the total error function 1618 calculates the total error of the initial model 20 20 according to, for example, equation 16 as described above. In some embodiments, it can
The total model error includes the individual errors combined to the model's prediction error for each data point in the total data set. Accordingly, the error function 1612 can also compute model errors according to, for example, Equation 5 as described above.
180476
-Ill-
In some embodiments, the model errors are used to determine a data record selection vector 01 in accordance with, for example, Equation 6 as described above. In some embodiments, the data record selection vector may include a binary classification based on the percentage of each model error for each data record in the model error distribution. In some embodiments, the data record selection vector includes
<p dir="rtl">5 On a threshold value of the percentage, where when the percentage is greater than the threshold value of the percentage, the data records are classified as outliers, and when it is equal to or less than that, the data records are classified as non-outliers. According to an illustrative model, the error function 1612 calculates a new data record selection vector (0 according to, for example, equation 6 as shown above to define the outlier data set 1617 and the non-outlier data set 1616. According to the model</p>
10 Illustratively, the data record selection module 1614 calculates which non-anomalous data records to include in a model calculation by selecting only records for which the record selection vector is equal to 1, according to, for example, Equation 7 as shown above.
Next, in accordance with an illustrative embodiment, the model 1608 calculates with the latest coefficients 1602 the new predicted values 1620 and the model coefficients 1602 from the data records identified by 1616 DOBR.
<p dir="rtl">15 According to, for example, equation 8 as shown above.</p>
Next, according to an illustrative model, the model 1608 computes, using the new model parameters, new predicted values 1620 for the complete data set. This step reproduces the calculation of the predicted values 1620 for the DOBR-selected records in the formal steps, but in practice the new model can be applied to the DOBR-selected records according to, for example, equation 9 as shown 20 above. Next, according to an illustrative model, the total error function 1618 computes the error
The total for the model according to, for example, equation 10 as described above.
Next, according to an illustrative embodiment, convergence test 1624 tests the convergence of the model according to, for example, equation 11 described above using convergence criteria 1622 (β), for example, 0.5%. In some embodiments, convergence test 1624 may terminate the iterative process if, for example,
180476
-112-
For example, the percentage error is less than, say, 0.5%. Otherwise, the process might return to the initial data set 1604.
In some embodiments, the anomaly analysis model 1609 may also use existing coefficients to calculate new predicted outliers 1621 and model outlier coefficients from the data set.
<p dir="rtl">5 1617 Anomaly according to, for example, Equation 8 as shown above. In some embodiments, similar to model 1608, the anomaly analysis model 1609 may be updated at each iterative step in reducing the anomaly dynamic bias. In some embodiments, the anomaly analysis model 1609 may be trained after all iterative steps in reducing the anomaly dynamic bias have been completed and convergence has been performed on the convergence criteria 1622 of model 1608. The anomaly analysis model may then be trained</p>
<p dir="rtl">10 Anomalous 1609 vs. anomalous data records for modeling bias-inducing anomalous values.</p>
In some embodiments, the anomaly analysis model 1609 may, for example, include a machine learning model suitable for modeling the anomaly data records, such as a regression model or a classifier model. For example, the anomaly analysis model 1609 may include, for example, decision trees, random forests, Naïve Bayes, K-nearest neighbors, support vector machines, or neutral networks.
<p dir="rtl">15 (convolutional neutral network and/or recurrent neutral network), or any other model or any combination thereof. In some embodiments, by training the anomaly analysis model 1609 using anomaly data records, the anomaly analysis model 1609 can ingest new data records to determine the likelihood of anomalous behavior. For example, extreme weather events can be predicted based on weather condition inputs provided to the anomaly analysis model 1609 trained on anomalous weather events. Based on</p>
<p dir="rtl">20 Accordingly, the anomaly analysis model 1609 may include a binary classifier model to classify data records as either probable anomalies or unlikely anomalies based on a predicted probability value. In some embodiments, where the predicted probability value exceeds a threshold probability value, relevant data records may be classified as probable anomalies. Such predictions may be used to inform predictions by model 1608 or other analyses based on the relevant data record.</p>
18476
-113-
In some embodiments, instead of a machine learning model, the anomaly analysis model 1609 may include a statistical model to characterize, for example, the frequency of outliers under certain conditions, the ratio of the frequency of outliers to the frequency of non-outliers under certain conditions, or other specification. In some embodiments, the frequencies and/or ratios may be based on, for example, the average values of
<p dir="rtl">5 Data records for specified conditions, or average values of data records for specified conditions, or other statistical compilation of data records under specified conditions.</p>
For example, the anomalous data records 1617 may be clustered according to the clustering model of the anomaly analysis model 1609, such as k-means clustering, distribution modeling (e.g., Bayesian distributions, mixture modeling, Gaussian modeling, etc.) or any other clustering analysis 10 or any combination thereof. As a result, the anomaly analysis model 1609 may cluster the anomaly records
1617 Anomalous data together according to similarity for use in, for example, cause analysis.
Radical or other analyses or any combination thereof.
DOBR for grid power optimization
Figures 17a to 17c illustrate the reduction of the anomalous dynamic bias for grid power demand forecasting.
<p dir="rtl">15 and improving power supplies according to an illustrative methodology according to one or more embodiments of the present disclosure.</p>
In some embodiments, one or more models may be trained to predict an output based on a given input 1706 x. In some embodiments, DOBR, such as the DOBR training engine 1505 and the filter 1501b described above, provides an iterative process for removing anomalous records subject to a predetermined criterion. This criterion is a user-specified error tolerance value expressed as a percentage. This criterion refers to the amount of
<p dir="rtl">20 The error that users are willing to accept in the model based on their insights and other analysis results.</p>
which will be described later in this discussion. A value of 100% indicates that all errors are accepted and no records will be removed in the DOBR process. If 0% is chosen, all records will be deleted. In general, error acceptance values in the range of 80 to 95% have been observed for industrial applications.
18476
-114-
In some embodiments, as described, reducing bias through reduced recursive and dynamic outliers in training machine learning models can provide efficient and robust training for more accurate machine learning models. In some embodiments, in addition to modeling a low outlier dataset, machine learning models can be applied in addition to other analytical models to the outlier dataset.
<p dir="rtl">5 This modeling of anomalous dataset can lead to the analysis of abnormal situations such as extreme events, external factors, anomalous factors and the root causes of such abnormal situations.</p>
Referring to Figure 17a, in some embodiments, a grid power demand model 1708 may be trained to predict grid power demand to optimize power supply and storage. Excess power supply may not be used.
<p dir="rtl">10 The electricity produced by the power generation facility, which leads to wastage of materials, resources and money required to supply the power. However, shortages in the supply of electrical energy can have serious consequences including blackouts and power outages that may be limited to a specific area or may be more widespread depending on the degree of shortage. Hence, the grid demand model 1708 is usefully trained to predict power demand more accurately which can provide improvements for power supply management and optimization.</p>
<p dir="rtl">15 To improve resource efficiency and reduce power outages.</p>
Accordingly, in some embodiments, the DOBR model training process, for example, by the DOBR training engine 1501 described above, may be provided with grid power demand training data 1704 to train the grid power demand model 1708 without anomalous bias. In some embodiments, the training data may include historical power data records, where each record contains a variable
20 Independent 1705 and target output variable 1706.
In some embodiments, the independent variable 1706 may contain grid condition data, such as time of day, day of week, day of month, month of year, season, weather, location, population density, among other electric power demand data and grid condition data. In some embodiments, the target output variable 1706 for each data record may contain, for example,
<p dir="rtl">25 A request for electrical power to the grid during a specific period or at a specific time. In some embodiments, this may be</p>
18476
-115-
The specified period or time contains, for example, the instantaneous date and time, a period of day, for example, morning, afternoon, night, two hours of the day, three hours of the day, four hours of the day, six hours of the day, eight hours of the day, twelve hours of the day, a day of the week, a day of the month, a month of the year, or any other period for evaluating the grid power demand.
<p dir="rtl">5 In some embodiments, the DOBR anomaly analysis may include a prior analysis where an error acceptance criterion (∝)1702 is chosen, such as ∝ = 80%. In some embodiments, the error acceptance criterion (∝)C may be determined according to, for example, Equation 1 as described above. In some embodiments, while other functional relationships may be used to adjust (∝)C, the percentile function is an intuitive guide to understanding why the model includes or excludes certain data records, such as Equation 2 as described above.</p>
<p dir="rtl">10 As shown above. Given the iterative DOBR procedure, in one embodiment, the convergence criterion can be specified.</p>
1724, for example, by 0.5%.
In some embodiments, each record of grid power demand training data 1704 may be provided to the grid demand model 1708 to generate a predicted output variable 1710 for each independent variable 1705. In some embodiments, the target output variable 1706 and the predicted output variable 1710 may have
<p dir="rtl">15 A network demand level, such as kilowatt (kW), gigawatt (GW), watt current (TW) or other unit of electrical power. Accordingly, to learn and predict the output according to the network condition data of the independent variable 1705, the network demand model 1708 may use an appropriate regression machine learning model. For example, the network demand model 1708 may contain, for example, Ridge regression, Lasso regression, decision tree, random forest, nearest neighbor,</p>
<p dir="rtl">20 K, support vector machine, neutral network (neutral recurrent network), or any other suitable regression model or any combination thereof. In some embodiments, DOBR may be implemented to reduce bias in the training of the network demand model 1708 M to more accurately predict future network demand levels without anomalous bias.</p>
In some embodiments, the network request model 1708 M is implemented by a model engine, including, for example, a processing device, memory, and/or storage device. Depending on the embodiments, the methodology calculates
18476
-116-
Illustrative model coefficients, (0)1 1702 and model estimates 24% 1710 for all records applying the network demand model 1708 M, to the full input data set (1704 according to, for example, Equation 3 as shown above.
Then, according to an illustrative model, the total error function 1718 calculates the total error of the initial model.
<p dir="rtl">5 20 According to, for example, Equation 17 as described above. In some embodiments, the total model error may include a model prediction error that aggregates the individual errors of the predicted network demand level compared to the target network demand level for the target output variable 1706 for each independent variable 1705. Accordingly, the error function 1712 may also compute model errors according to, for example, Equation 5 as described above.</p>
<p dir="rtl">10 In some embodiments, the model errors are used to determine a data record selection vector for * according to, for example, Equation 6 as described above. In some embodiments, the data record selection vector may have a binary classification based on the percentage of each model error for each data record in a distribution of model errors. In some embodiments, the data record selection vector includes a threshold value for the percentage, such that when the percentage is above that threshold value the classification is</p>
<p dir="rtl">15 Data records are classified as outliers, and when the ratio is equal to or less than that threshold value, data records are classified as non-outliers. According to an illustrative embodiment, the error function 1712 calculates a new data record selection vector (04) according to, for example, equation 6 as shown above to select the outlier data set 1717 and the non-outlier data set 1716. According to an illustrative embodiment, the data record selection module 1714 calculates the non-outlier data records to be included.</p>
<p dir="rtl">20 In the model calculation by selecting only records for which the record selection vector is equal to 1, according to, for example, Equation 7 as shown above.</p>
Next, according to an illustrative embodiment, the network demand model 1708 with the latest transactions 1702 calculates the new predicted network demand values 1720 and the model transactions 1702 from the data records identified by 1716 DOBR according to, for example, Equation 8 as shown above.
180476
-117-
Next, according to an illustrative embodiment, the network demand model 1708 calculates, using the new model coefficients, the new network demand values 1720 for the complete data set. This step reproduces the calculation of the new network demand values 1720 for DOBR-selected records in the formal steps, but in practice the new model can be applied to DOBR-selected records only according to, for example, Equation 9 as shown above. Then, according to an illustrative embodiment, the total error function 1718 calculates the total error of the model according to, for example, equation 10 as described above.
Next, according to an illustrative embodiment, convergence test 1724 tests the convergence of the model according to, for example, equation 11 described above using convergence criteria 1722 (β), for example, 10% 0.5. In some embodiments, convergence test 1724 may terminate the iterative process if,
For example, the error percentage is less than, say, 0.5%. Otherwise, the process can return to the initial data set 1704.
In some embodiments, to facilitate the identification of the source of power supply, the risk of extreme power requirements due to external factors may be identified by analyzing an anomalous data set resulting from the end of the process.
<p dir="rtl">15 DOBR. In some embodiments, as shown in Figure 17C, an extreme network demand model 1709 may be trained on an anomalous network demand dataset 1717 selected by the data record selection module 1714. In some embodiments, an extreme network demand model 1709 is trained to ingest an independent variable 1705 of the anomalous dataset 1717 and predict a risk 1721 for an extreme network demand event. For example, in some embodiments, certain conditions may be associated with an increased risk of 20 anomalous data that constitutes an abnormally high or abnormally low network demand as determined by the anomalous data record selection module 1714.</p>
In some embodiments, the extreme network demand model 1709 may use the target variable 1706 of the training dataset 1704 to determine the error in the predicted hazard 1721 and the updated model parameters of the extreme network demand model 1709. Hence, the extreme network demand model 25 1709 may be trained to predict a hazard score for an extreme network demand level.
18476
-118-
In some embodiments, referring to Figure 17B, a new network status data record 1731 may be measured for a power network 1730. For example, in some embodiments, the network status data record 1731 may include, for example, a time of day, day of week, day of month, month of year, season, weather, location, population density, among other electric power demand data that
<p dir="rtl">5 Power grid 1730.</p>
In some embodiments, based on the model parameters resulting from the termination of the recursive DOBR process, the network demand model 1708 may forecast a future demand level 1732. For example, in some embodiments, the forecast may include, for example, a network demand level over an hour, two hours, three hours, four hours, six hours, eight hours, 12 hours, 24 hours, two days, one week, ten weeks, one month, or other forecast period. Accordingly, the network demand model 1708 may produce
The expected level of network demand in the future.
In some embodiments, the extreme grid demand model 1709 may also receive a new grid condition data record 1731 measured for the power grid 1730. In some embodiments, the extreme grid demand model 1709 may ingest the grid condition data record 1731 and produce a prediction of the risk of extreme grid demand 15 1734 , such as a probability of an extreme grid condition value occurring based on training on demand levels
Anomalous depending on network conditions.
In some embodiments, the power generation facility 1733 may receive the expected grid demand level 1732 and the risk of extreme grid demand 1734 to optimize power generation. In some embodiments, the power generation facility 1733 may dynamically scale up power generation and storage to compensate for the expected increase or decrease in demand. In some embodiments, the dynamic metering may include
An optimization function that minimizes excess power generation while minimizing the risk of power shortages. For example, the power generation facility 1733 may balance the cost of excess power against the frequency or extent of shortages, thereby ensuring adequate power generation without wasting resources. In some embodiments, the power generation facility 1733 may also adjust dynamic metering where the risk of extreme grid demand 1734 is high, for example, above 50%, above 60%, above 75% or other appropriate threshold risk.
18476
-119-
For example, the power generation facility 1733 may generate and store additional electrical energy buffers (e.g., using batteries or other energy storage mechanisms) where the risk of extreme demand is high. As a result, the power generation facility 1733 may improve grid power supply management to reduce the risk of power shortages while also reducing resource inefficiencies.
<p dir="rtl">5 DOBR for user interaction with recommended content</p>
Figures 18A and 18B illustrate a reduction in the anomalous dynamic bias of user interaction-enhanced content recommendation prediction according to an illustrative methodology in accordance with one or more embodiments of the present disclosure.
In some embodiments, one or more models may be trained to predict an output based on a given input 1706 x. In some embodiments, DOBR, such as the DOBR training engine 1501 and the filter 1501b, provides
<p dir="rtl">10 Described above, it is an iterative process for removing anomalous records that meet a pre-defined criterion. This condition is a user-specified error acceptance value expressed as a percentage. This condition indicates the amount of error that the user is willing to accept in the model based on the results of their analysis and other analysis results that will be described later in this discussion. A value of 100% indicates that all errors are accepted and no records will be removed in the DOBR process. If 0% is selected,</p>
<p dir="rtl">15 All records will be deleted. In general, error tolerance values in the range of 80 to 95% have been observed for industrial applications.</p>
In some embodiments, referring to Figure 18a, reducing bias by iteratively and dynamically reducing outliers in machine learning model training can provide efficient and robust training to obtain more accurate machine learning models. For example, in some embodiments, a prediction model may be trained
<p dir="rtl">20 Content 1808 to predict content recommendations and/or content placement for computer users and software applications. For example, Internet advertisements may be placed on a web page that the user is browsing based on the content of the advertisement, or media content may be recommended in a media streaming application. Content predictions may be trained accordingly to optimize the user's interaction with the content during a browsing session.</p>
18476
-120-
Then, the content prediction model 1808 is usefully trained to more accurately predict content and position recommendations to increase user engagement.
Accordingly, in some embodiments, the DOBR model training process, for example, the DOBR training engine 1501 described above, may be provided with user interaction training data 1804 to train
<p dir="rtl">5 Content prediction model 1808 without anomalous bias. In some embodiments, the training data may contain the characteristics of each user and the degree of interaction with the content that each user encountered, where each record has an independent variable 1805 and a target output variable 1806.</p>
In some embodiments, the independent variable 1806 may contain user characteristic data, for example, user data such as user characteristics, including, for example,
<p dir="rtl">10 Browser, location, age, other user characteristics or any combination thereof, and user engagement metrics such as click-through rates, repeat clicks, times spent on content, among other user engagement metrics or any combination thereof. In some embodiments, the target output variable 1806 for each data record may include content attributes, for example, the source of the content, the location of the content on the page, the screen area of the content, the type of content or the classification.</p>
<p dir="rtl">15 In some embodiments, the outlier analysis using DOBR may include a pre-analysis where an error acceptance criterion (∝)1802 is specified, such as ∝ = 80%. In some embodiments, the error acceptance criterion (∝)C may be specified according to, for example, Equation 1 as described above. In some embodiments, while other functional relationships may be used to adjust C(α), the percentile function is an intuitive guide to understanding why the model includes or excludes certain data records, such as Equation 2</p>
<p dir="rtl">20 As shown above, since the DOBR procedure is iterative, in one embodiment, a</p>
The 1824 convergence standard, for example, is 0.5%.
In some embodiments, each user interaction training data record 1804 may be provided to the content prediction model 1808 to generate a predicted output variable 1810 for each independent variable 1805. In some embodiments, the target output variable 1806 and the predicted output variable 1806 may contain
18476
-121-
1810 On the content characteristics to determine the content to be displayed to the user such as the content source, the content location on the page, the content screen area, the content type or classification. Accordingly, to learn and predict the output according to the user characteristics data of the independent variable 1805, the content prediction model 1808 may use an appropriate classifier machine learning model, for example
<p dir="rtl">5 Example, multi-label classifier. For example, the content prediction model 1808 may include, for example, collaborative filtering, logistic regression, decision tree, random forest, K-nearest neighbor, support vector machine, neutral network (e.g., convolutional neutral network), or any other suitable classifier model or any combination thereof. In some embodiments, DOBR may be implemented to reduce bias in the training of the content prediction model 1808 M to more accurately predict</p>
<p dir="rtl">10 At future network demand levels without anomalous bias.</p>
In some embodiments, the content prediction model 1808M is implemented by a model engine, including, for example, a processing device, memory, and/or storage device. According to one embodiment, the illustrative methodology calculates model coefficients, (6)1 1802 and model estimates, (32)1810 for all records applying the content prediction model 1808M, for the entire input data set.
<p dir="rtl">15 (-3(+] According to, for example, equation 3 as shown above.</p>
Then, according to an illustrative embodiment, the total error function 1818 computes the total error of the initial model 20 according to, for example, equation 18 as shown above. In some embodiments, the total model error may include a model prediction error that aggregates the individual errors of the predicted population characteristics compared to the target population characteristics of the target output variable 20 1806 for each independent variable 1805. Accordingly, the error function 1812 may also compute
Model errors according to, for example, Equation 5 as shown above.
For example, in some embodiments, the predicted output variable 1810 may be compared to the target variable 1806 to assess the prediction error. In some embodiments, the error may be influenced by an optimizer that uses a loss function to maximize user engagement metrics, e.g., rates of
<p dir="rtl">25 Clicks, click frequency, times spent on content, among other user engagement metrics.</p>
180476
-122-
Other or any combination thereof. Accordingly, the error is based on the difference between the predicted output variable 1810 and the target variable 1806 according to the levels of user interaction to update the parameters of the content prediction model 1808.
In some embodiments, model errors are used to determine a data record selection vector (according to,
<p dir="rtl">5 For example, for Equation 6 as described above. In some embodiments, the data record selection vector may include a binary classification based on the percentage of each model error for each data record in the model error distribution. In some embodiments, the data record selection vector includes a threshold value for the percentage, such that when the percentage is greater than that threshold value, data records are classified as outliers, and when the percentage is equal to or less than that threshold value, they are</p>
<p dir="rtl">10 Classifying data records as non-anomous values. According to an illustrative embodiment, the error function 1812 computes a new data record selection vector (0) according to, for example, equation 6 as shown above to select the anomalous data set 1817 and the non-anomalous data set 1816. According to an illustrative embodiment, the data record selection module 1814 computes the non-anomalous data records to be included in the model calculation by selecting only records for which the record selection vector is equal to</p>
<p dir="rtl">15 1, According to, for example, equation 7 as shown above.</p>
Next, according to an illustrative embodiment, the content prediction model 1808 with the latest coefficients 1802 calculates the predicted new content properties 1820 and the model coefficients 1802 from the specified data records 1816 DOBR according to, for example, equation 8 as shown above.
Then, according to an illustrative model, the content prediction model 1808 is calculated using coefficients.
20 New model, new content properties 1820 for the entire dataset. This step recalculates the new content properties 1820 for the records selected by DOBR in the formal steps, but in practice the new model can be applied to the records removed by DOBR according to, for example, equation 9 as shown above. Then, according to an illustrative model, the total error function 1818 calculates the total error of the model according to, for example, equation
25 10 As described above.
180476
-123-
Next, according to an illustrative embodiment, the convergence test 1824 tests the model for convergence according to, for example, equation 11 described above using the convergence criteria 1822 (β), for example, 0.5%. In some embodiments, the convergence test 1824 may terminate the iterative process if, for example, the error rate is less than, for example, 0.5%. Otherwise, the process may return to the initial data set 1804.
In some embodiments, referring to Figure 18B, new user properties 1831 of a user viewing content on a user computing device 1830 may contain. For example, in some embodiments, user properties 1831 may contain, for example, a browser, software application, device identifier, location, age or other user properties or any combination 10 thereof that distinguishes the user on a user computing device 1830.
In some embodiments, based on the model parameters resulting from the completion of the iterative DOBR process, the content prediction model 1808 may predict the content characteristics 1832 of the content to be displayed to the user to increase user engagement. For example, in some embodiments, the prediction may include, for example, the source of the content, the location of the content on the page, the screen area of the content, the type 15 of the content, the taxonomy, or any combination thereof. Accordingly, the prediction model may produce
With content 1808 content recommendations and placement to increase engagement.
In some embodiments, the user computing device 1820 may receive content selected according to the characteristics of the content 1832 to display the content to the user. Accordingly, the user computing device 1830 may automatically receive dynamically selected content to maximize user engagement, e.g., improve advertising revenue, more precise advertising subject matter, media (such as
Music, video, images, social media content, etc.) that most closely match user behavior, etc. Accordingly, the DOBR process can improve the content prediction model 1808 to provide content according to the content characteristics 1832 with low bias due to anomalous behavior.
18476
-124-
In some embodiments, and optionally, in a combination of any of the embodiments described above or below, the DOBR explanatory machine learning model may be based at least in part on Monte Carlo-type computational algorithms (e.g., Solovay–Strassen-type algorithms, Baillie–PSW-type algorithms, Miller–Rabin-type algorithms, and/or Schreier-type algorithms).
<p dir="rtl">5 SIMS(which may take into account historical quality data for the required non-anomalous data. In some embodiments, and optionally, in combination with any model described above or below, the DOBR explanatory machine learning model may be continuously trained, for example, by applying at least one machine learning technique (for example, but not limited to, decision trees, boosting, support vector machines, neutral networks, nearest neighbor algorithms, Naive Bayes, bagging,</p>
<p dir="rtl">10 Random forests, etc.) on collected and/or aggregated sensor data (e.g., various types of visual data about the appearance of the environment and/or the physical/visual appearance of the goods). In some embodiments, and optionally, in combination with any embodiment described above or below, the illustrative neutral network technology can be one of, but not limited to, a feedforward neutral network, a radial basis function network, a recurrent neutral network, a convolutional network (e.g.</p>
<p dir="rtl">15 For example, U-net( or any other suitable network. In some embodiments, and optionally, in a combination of any embodiment described above or below, an illustrative implementation of the neutral network may be performed as follows:</p>
<p dir="rtl">1( Determine the neutral network structure/model,</p>
<p dir="rtl">2(Transfer input data to the neutral network model,</p>
<p dir="rtl">3( Train the illustrative model gradually,</p>
<p dir="rtl">20 4( Determine the accuracy for a given number of time steps,</p>
<p dir="rtl">5( Apply the illustrative training model to process the newly received input data,</p>
<p dir="rtl">6(Optionally and in parallel, continue training the trained model with a pre-defined periodicity.</p>
18476
-125-
In some embodiments, optionally, in combination with any model described above or below, the trained neutral network model can specify a neutral network by a neutral network structure, a series of activation functions, and connection densities. For example, a neutral network structure can include a configuration of neutral network nodes and connections between such nodes. In some embodiments, optionally, in combination with any model described above or below, the trained neutral network model can also specify
The illustrative function may include other parameters, including but not limited to bias values/functions and/or summation functions. For example, the activation function for a node may be a step function, a sine function, a continuous or piecewise linear function, a sigmoid function, a hyperbolic tangent function, or some other type of mathematical function that represents a threshold value at which the node is activated. In some
<p dir="rtl">10 In some embodiments, optionally, in combination with any model described above or below, the illustrative aggregation function may be a mathematical function that integrates (e.g., sum, product, etc.) the input signals of the node. In some embodiments, optionally, in combination with any model described above or below, the output of the illustrative aggregation function may be used as an input to the illustrative activation function. In some embodiments, optionally, in combination with any model described above or below, the</p>
<p dir="rtl">15 A bias is a constant value or function that can be used by an aggregation function and/or an activation function to make a node more or less likely to be activated.</p>
In some embodiments, and optionally, in combination with any embodiment described above or below, illustrative connection data for each connection in the illustrative neutral network may include at least one node pair or connection density. For example, if the illustrative neutral network includes
<p dir="rtl">20 Connection from node N1 to node N2, the connection data for this connection may include:</p>
Node pair <N1, N2<. In some embodiments, and optionally, in combination with any embodiment described above or below, the connection density may be a scalar quantity that affects whether and/or how the output of N1 is modified before being input at N2. In the example of a recurrent network, a node may have a connection itself (e.g., the connection data may include the node pair <N1, N1<).
18476
-126-
In some embodiments, and optionally, in combination with any model described above or below, the trained neutral network explanatory model may also include a type identifier (ID) and relevance data. For example, each type identifier may indicate which set of model types (e.g., cargo loss classes) the model is classified into. For example, the relevance data may indicate
<p dir="rtl">5 Fit refers to how well the trained neutral network model models the sensory input dataset. For example, fit data can include a fit value that is determined based on the evaluation of the fit function with respect to the model. For example, the fit function may be an objective function based on the frequency and/or magnitude of errors resulting from testing the trained neutral network model on the sensory input dataset. As a simple example, it is assumed that</p>
<p dir="rtl">10 The input sensor dataset has ten rows, the input sensor dataset has two columns labeled A and B, and the trained neutral network model produces a predicted value for B given an input value of A. In this example, the test of the trained neutral network model could have each of the ten values for A from the input sensor dataset be input, and the predicted values for B be compared to the actual values for B from the sensor dataset.</p>
<p dir="rtl">15 The input, and determine whether there is a difference between the predicted value and the actual value of (b) and/or the amount of the difference. To illustrate, if a given neutral network correctly predicts the value of (b) for nine of the ten rows, the explanatory fit function might assign the corresponding model a fit value of 10/9 = 0.9. It should be understood that the previous example is for illustration only and should not be considered restrictive. In some embodiments, the explanatory fit function can be based on factors unrelated to the frequency of error or the rate of</p>
<p dir="rtl">20 Error, such as number of input nodes, layers of nodes, hidden layers, connections, computational complexity, etc.</p>
In some embodiments, and optionally, in combination with any embodiment described above or below, the present disclosure may utilize multiple aspects of at least one of:
The American patent, serial number 8195484, is titled “Insurance product
"rating system and method 25
18476
-127-
The American patent, serial number 8548833, is titled “Insurance product
rating system and method
The American patent, serial number 8554588, is titled “Insurance product
rating system and method
<p dir="rtl">5 The American patent, serial number 8554589, is titled “Insurance product</p>
rating system and method
The American patent, serial number 8595036, is titled “Insurance product
rating system and method
The American patent, serial number 8676610, is titled “Insurance product
"rating system and method 10
The American patent, serial number 8719059, is titled “Insurance product
rating system and method
The American patent, serial number 8812331, entitled “Insurance product
rating and credit enhancement system and method for insuring project “savings 15
Some aspects of the present disclosure will now be described at least by reference to the following numbered items:
Item 1. Method including:
receiving, through at least one processor, a training data set of target variables representing at least one activity-related attribute of at least one user activity;
<p dir="rtl">20 Received - by at least one processor - at least one bias coefficient that was used to identify one or more outliers;</p>
18476
-128-
Specify - by at least one processor - a set of model parameters for a machine learning model that includes:
(1) Applying - by at least one processor - a machine learning model containing a set of initial model parameters to a training data set to determine a set of predicted model values;
<p dir="rtl">5 )2( Generating - by at least one processor - a set of errors from data element errors.</p>
The method of comparing the set of predicted model values to the corresponding actual values of the training data set;
<p dir="rtl">) 3 ( Generating - by at least one processor - a data selection vector to identify non-anomalous target variables, based at least in part on the set of data element errors and at least one bias criterion;</p>
<p dir="rtl">10 )4( Using - by at least one processor - a data selection vector on a data set.</p>
Training to generate a non-anomalous dataset;
<p dir="rtl">(5) Determine - by at least one processor - a set of updated model parameters for the machine learning model based on the non-anomalous data set; and</p>
<p dir="rtl">)6( Repeat - by one processor at least - steps )1(-)5( Repeat - by processor</p>
<p dir="rtl">15 At least one iteration of steps (1)-5 until the at least one control performance termination criterion is met to obtain the model parameter set of the at least one machine learning model in the form of the updated model parameters, where each iteration regenerates the set of predicted values, the set of errors, the data selection vector, and the non-anomalous data set using the updated model parameter set as the initial model parameter set;</p>
<p dir="rtl">20 Training—by at least one processor—a set of classifier model parameters for an anomaly classifier machine learning model to obtain a trained anomaly classifier machine learning model that is initialized to select at least one anomalous data item based at least in part on the training data set and the data selection vector;</p>
18476
-129-
Applying - by at least one processor - the trained machine learning model of the anomaly classifier to a dataset of activity-related data for at least one user activity to determine:
<p dir="rtl">1( A set of anomalous data related to activity in a data set The data is related to activity, and</p>
<p dir="rtl">2( A set of non-anomalous data related to the activity in the data set of the data related to</p>
<p dir="rtl">5 with activity; and</p>
Applying - by at least one processor - a machine learning model to a set of non-anomalous activity-related data elements to predict a characteristic of future activity relevant to at least one user's activity.
Article 2. A system comprising:
<p dir="rtl">10 At least one processor connected to a non-temporary computer-readable storage medium with software instructions stored therein, wherein the software instructions, when executed, cause at least one processor to perform the following steps:</p>
Receive a training dataset of target variables that represent an attribute associated with at least one activity for at least one user activity;
<p dir="rtl">15 Receive at least one bias criterion used to identify one or more outliers;</p>
Specify a set of model parameters for a machine learning model, including:
<p dir="rtl">(1) Apply a machine learning model that includes a set of initial model parameters to the training data set to determine a set of predicted model values;</p>
<p dir="rtl">)2( Generate an error set of data element errors by comparing a set of model values.</p>
<p dir="rtl">20 predicted with the actual values corresponding to the training data set;</p>
<p dir="rtl">(3) Generate a data selection vector to identify non-anomalous target variables based at least partially on the set of data element errors and at least one bias criterion;</p>
18476
-130-
<p dir="rtl">)4) Use the data selection vector on the training dataset to generate a non-outlier dataset;</p>
<p dir="rtl">)5) Determine a set of updated model parameters for the machine learning model based on the non-anomalous dataset; and</p>
<p dir="rtl">5 )6( Repeat steps (1)-(5) with a specified frequency until the criterion for terminating one control tool is met.</p>
To obtain the model parameter set of the machine learning model as the updated model parameters, each iteration regenerates the predicted value set, error set, data selection vector, and non-outlier data set using the updated model parameter set as the initial model parameter set;
<p dir="rtl">10 Training a set of classifier model parameters for an anomaly classifier machine learning model to obtain a trained anomaly classifier machine learning model that is initialized to identify at least one anomaly data item based at least partially on the training data set and the data selection vector;</p>
Apply the trained machine learning model of the anomaly classifier to a dataset of activity-related data for at least one user activity to determine:
<p dir="rtl">15 1(a set of anomalous data related to activity in a data set of activity-related data, and</p>
<p dir="rtl">2(a) a set of non-anomalous data relating to the activity in the activity data set; and</p>
Apply a machine learning model to a set of non-anomalous activity-related data items to predict the attribute of future activity relevant to at least one user's activity.
<p dir="rtl">20 Item 3. The systems and methods of Items 1 and/or 2, which also include:</p>
Applying - by at least one processor - a data selection vector to the training data set to identify an outlier training data set;
18476
-131-
Train—by at least one processor—using the outlier training dataset, at least one outlier-specific model parameter of the at least one outlier-specific machine learning model to predict outlier data values; and
Use - by at least one processor - an outlier machine learning model to predict 5 activity-related outlier values for a set of activity-related outliers.
Item 4. The systems and methods of Items 1 and/or 2, which also include:
Training - by at least one processor - using a training dataset, general model parameters of a general machine learning model to predict data values;
Use - with at least one processor - a general machine learning model to predict outliers in data.
<p dir="rtl">10 activity-related anomaly data set; and</p>
Using - by at least one processor - a general machine learning model to predict activity-related data values.
Item 5. The systems and methods of Items 1 and/or 2, which also include:
Apply - by at least one processor - a data selection vector to the training data set.
<p dir="rtl">15 To identify an outlier training dataset;</p>
Training - by at least one processor - using the outlier training dataset, outlier-specific model parameters of an outlier-specific machine learning model to predict outlier data values;
Training - by at least one processor - using the training dataset, general model parameters of a general machine learning model to predict data values;
<p dir="rtl">20 Using - by at least one processor - an anomaly machine learning model to predict anomaly values for activity-related data for an anomaly data set; and</p>
18476
-132-
Using - by at least one processor - an anomaly-specific machine learning model to predict activity-related data values.
Item 6. The systems and methods of Items 1 and/or 2, which also include:
Training - with at least one processor - using the training dataset, model parameters
<p dir="rtl">5 General machine learning model for predicting data values;</p>
Using - by at least one processor - a general machine learning model to predict activity-related data values for an activity-related data set;
Using - by at least one processor - an anomaly classifier machine learning model to identify anomalous activity-related data values for activity-related data values; and
<p dir="rtl">10 Remove - by at least one processor - anomalous data values related to the activity.</p>
Item 7. Systems and methods of Items 1 and/or 2, wherein the training dataset includes at least one activity-related characteristic of concrete compressive strength as a function of concrete composition and exposure to concrete curing.
Item 8. Systems and methods of Items 1 and/or 2, where the training dataset includes a single feature
<p dir="rtl">15 At least activity-related energy use data as a function of household environmental and lighting conditions.</p>
Item 9. The systems and methods of Items 1 and/or 2, which also include:
Receiving - by at least one processor - an API request to generate a prediction for at least one data item; and
<p dir="rtl">20 Create an instance—with at least one processor—of at least one cloud computing resource to schedule the execution of the machine learning model;</p>
18476
-133-
Using - by at least one processor - according to the execution schedule, a machine learning model to predict the value of
at least one activity-related data element for at least one data element; and
Returning - by at least one processor - at least one activity-related data item value to a computing device associated with an API request.
<p dir="rtl">5 Item 10. Systems and methods of Items 1 and/or 2, where the training dataset includes a single feature</p>
At least activity-related 3D patient images of a medical dataset; and
Where the machine learning model is configured to predict activity-related data values that include two or
More than physical display parameters based on medical dataset.
Item 11. Systems and methods of Items 1 and/or 2, where the training dataset includes a single feature
<p dir="rtl">10 At least related to the activity of the simulated control results and the command of the electronic machine; and</p>
Where the machine learning model is configured to predict activity-related data values that include commands.
Electronic control.
Item 12. The systems and methods of Items 1 and/or 2, which also include:
Dividing - by at least one processor - a set of activity-related data into a set of
<p dir="rtl">15 Subsets of activity-related data;</p>
Determine - by at least one processor - a consistent set model for each subset of
Activity related data;
Where a machine learning model includes a consistent set of models;
Where each harmonic set model contains a random combination of models from the harmonic set.
<p dir="rtl">20 From the models;</p>
Using - by at least one processor - each model of a coordinated set separately to predict values
Activity data for a consistent group;
18476
-134-
Determine - by at least one processor - an error for each consistent set model based on the activity-related data values of the consistent set and known values; and
Selection - by at least one processor - of the consistent set model that has the best performance based on the least error.
<p dir="rtl">5 The publications described herein are incorporated in their entirety by reference. While one or more embodiments of the present disclosure are described, it is understood that such embodiments are only illustrative and not limiting, and that many modifications may become apparent to one of ordinary skill in the art, including various embodiments of inventive methodologies, where the underlying software systems and inventive devices described herein are used in any combination with each other. Furthermore, the various steps may be performed</p>
<p dir="rtl">10 In any desired order (any desired steps can be added and/or any desired steps can be removed).</p>
18476
-135-
Contents4
1 sheet
Sheet 1
51 members in 14 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 201962902074 | United States of America | P | |
| 17025889 | United States of America | – | |
| 202017025889 | United States of America | A | |
| 2021022861 | United States of America | W |
Members51
| Document | Office | Kind | |
|---|---|---|---|
| CA3154671A1 | Canada | A1 | |
| WO2021055847A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2021110313A1 | United States of America | A1 | |
| CA3195894A1 | Canada | A1 | |
| US2022092346A1 | United States of America | A1 | |
| WO2022060411A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US11288602B2 | United States of America | B2 | |
| US11328177B2 | United States of America | B2 | |
| GB202204238D0 | United Kingdom | D0 | |
| KR20220066924A | Republic of Korea | A | |
| CN114556382A | China | A | |
| EP4022532A1 | European Patent Office (EPO) | A1 | |
| GB2603358A | United Kingdom | A | |
| US2022277232A1 | United States of America | A1 | |
| BR112022005003A2 | Brazil | A2 | |
| US2022284235A1 | United States of America | A1 | |
| JP2022548654A | Japan | A | |
| US11599740B2 | United States of America | B2 | |
| US11615348B2 | United States of America | B2 | |
| NO20230419A1 | Norway | A1 | |
| AU2021343372A1 | Australia | A1 | |
| KR20230070272A | Republic of Korea | A | |
| GB202305640D0 | United Kingdom | D0 | |
| US2023169153A1 | United States of America | A1 | |
| MX2023003217A | Mexico | A | |
| DE112021004908T5 | Germany | T5 | |
| GB2614849A | United Kingdom | A | |
| EP4214652A1 | European Patent Office (EPO) | A1 | |
| CN116569189A | China | A | |
| GB2603358B | United Kingdom | B | |
| GB2617045A | United Kingdom | A | |
| JP7399269B2 | Japan | B2 | |
| US11914680B2 | United States of America | B2 | |
| JP2024026276A | Japan | A | |
| GB202402945D0 | United Kingdom | D0 | |
| GB2617045B | United Kingdom | B | |
| GB2625937A | United Kingdom | A | |
| EP4414902A2 | European Patent Office (EPO) | A2 | |
| US2024311446A1 | United States of America | A1 | |
| AU2021343372B2 | Australia | B2 | |
| EP4414902A3 | European Patent Office (EPO) | A3 | |
| SA18476B1This record | Saudi Arabia | B1 | |
| SA523440223B1 | Saudi Arabia | B1 | |
| AU2024278534A1 | Australia | A1 | |
| GB2625937B | United Kingdom | B | |
| US12223018B2 | United States of America | B2 | |
| SA522431988B1 | Saudi Arabia | B1 | |
| GB202500423D0 | United Kingdom | D0 | |
| KR102778732B1 | Republic of Korea | B1 | |
| JP7654916B2 | Japan | B2 | |
| GB2639108A | United Kingdom | A |
Numbers
- Publication
- 18476
- Application
- 523440223
Titles2
- Arabic
- أنظمة قائمة على الحاسوب ومكونات حاسوبية وعناصر حاسوبية مهيأة لتخفيض الانحياز الديناميكي الشاذ في نماذج التعلم الآلي
- English
- COMPUTER-BASED SYSTEMS, COMPUTING COMPONENTS AND COMPUTING OBJECTS CONFIGURED TO IMPLEMENT DYNAMIC OUTLIER BIAS REDUCTION IN MACHINE LEARNING MODELS
Classification
- CPC, 11
- G06N20/00
- G06N20/20
- G06F17/18
- G06N3/08
- G06F9/4881
- G06F18/2433
- G06F18/214
- G06N5/01
- G06N7/01
- G06N3/09
- G06N3/0985
- IPC, 2
- G06N20 00
- G06N200 00